Benchmark: Predicting Real
Human Behavior
We test Mimiq against published randomized controlled trials and real-world survey data. Every ground truth number comes from a citable source. Here's what we found, and where we fall short.
How we test
We take published experiments where real human behavior was measured under controlled conditions. Then we run Mimiq's AI personas through the same scenarios and compare: did Mimiq predict the same outcome as the real experiment?
Find real experiments
We collect published A/B tests and surveys with verified outcomes. Sample sizes range from 4,000 to 2.4 million real participants.
Run Mimiq personas
For each experiment, we generate a matching cohort at the run’s configured sample size. Each persona independently evaluates the content and decides what to do.
Compare to reality
Did Mimiq pick the same winner? Are the predicted distributions close to the real ones? We measure directional accuracy and distribution distance.
The sycophancy problem
LLMs are trained to be helpful and agreeable. When role-playing as users, they default to positive responses, producing conversion predictions of 40-85% when real rates are 2-5%. This makes naive LLM persona approaches no better than a coin flip.
* Latest 23-test fast A/B run. Direction is useful; numeric lift forecasting is not yet reliable. See limitations.
A/B test results: Mimiq vs. reality
For each published experiment below, we show the real-world outcome and whether Mimiq correctly predicted the winner. The question is simple: given two variants, can Mimiq tell you which one will perform better?
Total real-world participants across all studies: 3,500,000+
This is not an overall accuracy estimate. In the broader 23-test fast suite—including difficult Upworthy headline tests—Mimiq reached about 78% winner-direction accuracy, while only about 17% of tests met the current effect-size realism threshold. Use it to rank and diagnose; validate forecasted lift with live behavior.
Counter-intuitive tests
These are the hardest tests: experiments where the "obvious" answer is wrong. Standard LLM personas consistently fail these because they default to agreeing that interventions help.
Distribution realism
The values below are a static calibration snapshot against industry reference rates, not a live accuracy guarantee. Re-run the versioned suite before citing a prediction externally.
SaaS
E-commerce
Financial Services
Business Services
Anti-sycophancy tests
Can the model say no? We test Mimiq against obviously bad pages, empty value propositions, and genuinely good pages to verify that rejection scales with quality.
A standard GPT persona told a scam page was "compelling" and predicted 60%+ engagement. Mimiq's pass criteria: scam pages must get <5% conversion and >80% bounce.
Survey benchmark coverage
The registry contains 50 distribution benchmarks across five domains. We report performance from versioned runs using total variation distance and correlation; an undefined aggregate “accuracy” percentage is intentionally not shown.
Tech Adoption
Consumer Behavior
Workplace Trends
Social Sentiment
Health & Lifestyle
Limitations
We believe in stating limitations honestly. Here's what doesn't work well yet.
Effect size magnitude is weak
HIGHMimiq often predicts the winning direction, but not by how much. Predicted effect sizes showed weak correlation with actual magnitudes (r = 0.04 on the curated 7-test suite), and only about 17% passed the broader suite’s realism threshold. Useful for ranking variants, not reliable for precise lift forecasting.
Selected cases are not an accuracy estimate
HIGHThe seven displayed cases all match direction, but they are a curated explanatory set. The broader 23-test fast run reached about 78% directional accuracy. A locked holdout and reproducible run manifest are required before making a generalized accuracy claim.
Domain coverage is uneven
MEDIUMPerformance is strongest on e-commerce, SaaS, and content engagement. Specialized B2B, medical, and government domains have higher variance in predictions.
Not using a custom model yet
MEDIUMAll results use a general-purpose LLM with Mimiq's proprietary behavioral system. A dedicated behavioral prediction model is in training and expected to improve both accuracy and calibration.
Headline A/B tests are hard
MEDIUMUpworthy-style headline tests show the largest direction and effect-size errors. The model systematically under-predicts the magnitude of headline effects and sometimes picks the wrong winner.
Some responses fail parsing
LOWA small percentage (typically 2-5%) of model responses cannot be parsed into valid actions. These are excluded from results, which adds noise and reduces effective sample size.
Custom model in development
The results above are the current behavioral-system baseline. We are developing a dedicated propensity model intended to improve accuracy and calibration; no improvement is assumed until it passes a locked holdout.
Anti-sycophancy training
The model learns to say "no" from real negative feedback: product reviews, complaints, and sycophancy research datasets.
Real outcome data
Ground truth from published RCTs and the benchmark suite documented on this page.
Rejection patterns
Teaching the model to produce realistic objections, not generic praise. The negative signal is the product.
References
Run your own benchmark.
Test your landing page, pricing, or campaign against a simulated audience.
Try Mimiq Free