ALPHACustom behavioral model in development

Benchmark: Predicting Real
Human Behavior

We test Mimiq against published randomized controlled trials and real-world survey data. Every ground truth number comes from a citable source. Here's what we found, and where we fall short.

7
A/B test benchmarks
from published RCTs
50
Survey benchmarks
5 domains
3.5M+
Ground truth participants
across all studies
139+
Data sources
Census, Pew, Gallup, OECD

How we test

We take published experiments where real human behavior was measured under controlled conditions. Then we run Mimiq's AI personas through the same scenarios and compare: did Mimiq predict the same outcome as the real experiment?

STEP 1

Find real experiments

We collect published A/B tests and surveys with verified outcomes. Sample sizes range from 4,000 to 2.4 million real participants.

STEP 2

Run Mimiq personas

For each experiment, we generate a matching cohort at the run’s configured sample size. Each persona independently evaluates the content and decides what to do.

STEP 3

Compare to reality

Did Mimiq pick the same winner? Are the predicted distributions close to the real ones? We measure directional accuracy and distribution distance.

The sycophancy problem

LLMs are trained to be helpful and agreeable. When role-playing as users, they default to positive responses, producing conversion predictions of 40-85% when real rates are 2-5%. This makes naive LLM persona approaches no better than a coin flip.

STANDARD LLM PERSONA
"This is a well-designed landing page with clear messaging and a compelling value proposition. I would definitely sign up."
~50%
A/B accuracy (random)
40-85%
Predicted conversion
MIMIQ PERSONA
"Looks clean but I already use Asana. Not switching tools mid-sprint for some new thing I've never heard of. Maybe later."
~78%
Winner direction · 23 tests*
17%
Effect-size realism pass*

* Latest 23-test fast A/B run. Direction is useful; numeric lift forecasting is not yet reliable. See limitations.

A/B test results: Mimiq vs. reality

For each published experiment below, we show the real-world outcome and whether Mimiq correctly predicted the winner. The question is simple: given two variants, can Mimiq tell you which one will perform better?

Total real-world participants across all studies: 3,500,000+

Curated published-case set
Mimiq picked the same winner in each of the seven cases displayed below.
7/7
selected case set

This is not an overall accuracy estimate. In the broader 23-test fast suite—including difficult Upworthy headline tests—Mimiq reached about 78% winner-direction accuracy, while only about 17% of tests met the current effect-size realism threshold. Use it to rank and diagnose; validate forecasted lift with live behavior.

Hillstrom Email RCT (Mens)

GOLD
MineThatData, 2008 · n = 16,761
Email marketingCORRECT
No email
10.6%
Mens email
WINNER
18.1%+71%
Real winner: Mens email · Mimiq predicted the same winner

Hillstrom Email RCT (Womens)

GOLD
MineThatData, 2008 · n = 16,919
Email marketingCORRECT
No email
10.6%
Womens email
WINNER
15.3%+45%
Real winner: Womens email · Mimiq predicted the same winner

Cookie Cats Gate Retention

SILVER
Tactile Entertainment / Kaggle · n = 90,091
Mobile gamingCORRECT
Gate at level 30
WINNER
19.0%
Gate at level 40
18.2%-5%
Real winner: Gate at level 30 · Mimiq predicted the same winner

Udacity Landing Page

SILVER
Udacity, 2017 · n = 290,348
Landing pageCORRECT
Old landing page
WINNER
12.0%
New landing page
11.9%-1%
Real winner: Old landing page · Mimiq predicted the same winner

Upworthy Headline (Sitting/Health)

GOLD
Nature Scientific Data, 2021 · n = 4,136
Content engagementCORRECT
Practical framing
0.52%
Benefit framing
WINNER
1.4%+163%
Real winner: Benefit framing · Mimiq predicted the same winner

Upworthy Headline (Media/Women)

GOLD
Nature Scientific Data, 2021 · n = 10,454
Content engagementCORRECT
Vague framing
1.9%
Specific framing
WINNER
3.1%+62%
Real winner: Specific framing · Mimiq predicted the same winner

Price Anchoring Effect

SILVER
Tversky & Kahneman replications · n = 30,000
E-commerce pricingCORRECT
Low anchor ($99)
3.4%
High anchor ($199)
WINNER
5.2%+53%
Real winner: High anchor ($199) · Mimiq predicted the same winner

Counter-intuitive tests

These are the hardest tests: experiments where the "obvious" answer is wrong. Standard LLM personas consistently fail these because they default to agreeing that interventions help.

UK Organ Donor Registration

GOLD
UK Behavioural Insights Team, 2013 · n = 675,000
CORRECT
Control
2.3%
Treatment
3.6%

Combined reciprocity + loss framing outperformed all other message variants

UK Tax Debt Reminder Letters

GOLD
Hallsworth et al. / NBER, 2017 · n = 196,737
CORRECT
Control
33.9%
Treatment
38.8%

"You are in the very small minority" framing was best performer across 7 variants

Wikipedia Donation Social Proof

GOLD
Linek & Traxler, 2021 (J. Public Economics) · n = 2,387,700
CORRECT
Control
0.112%
Treatment
0.091%

Counter-intuitive: "Already 115,000 donated" REDUCED donations. Mimiq correctly predicted the backfire.

Distribution realism

The values below are a static calibration snapshot against industry reference rates, not a live accuracy guarantee. Re-run the versioned suite before citing a prediction externally.

Real-world
Mimiq prediction

SaaS

Unbounce 2024 (41K pages, 57M conversions)
Converted
3.8%
2.0%
Engaged
25.0%
18.0%
Bounced
71.2%
80.0%

E-commerce

Unbounce 2024 + Contentsquare 2025
Converted
5.5%
4.0%
Engaged
30.0%
24.0%
Bounced
64.5%
72.0%

Financial Services

Unbounce 2024
Converted
8.4%
6.0%
Engaged
22.0%
20.0%
Bounced
69.6%
74.0%

Business Services

Unbounce 2024
Converted
5.2%
4.0%
Engaged
28.0%
22.0%
Bounced
66.8%
74.0%

Anti-sycophancy tests

Can the model say no? We test Mimiq against obviously bad pages, empty value propositions, and genuinely good pages to verify that rejection scales with quality.

TEST
MAX CONVERSION
MIN BOUNCE
Obviously bad landing page
"Make money fast" scam page with fake testimonials
5%
80%
Empty value proposition
Generic page: "We help businesses do better things"
8%
70%
Well-designed page (Notion)
Real Notion-style page. Should convert reasonably.
30%
30%

A standard GPT persona told a scam page was "compelling" and predicted 60%+ engagement. Mimiq's pass criteria: scam pages must get <5% conversion and >80% bounce.

Survey benchmark coverage

The registry contains 50 distribution benchmarks across five domains. We report performance from versioned runs using total variation distance and correlation; an undefined aggregate “accuracy” percentage is intentionally not shown.

Tech Adoption

Pew Research, Stanford AI Index, Gallup
10 definitions

Consumer Behavior

McKinsey, Deloitte, PwC
10 definitions

Workplace Trends

Microsoft Work Trend Index, Gallup
10 definitions

Social Sentiment

Edelman Trust Barometer, Pew
10 definitions

Health & Lifestyle

Gallup, IPSOS, Deloitte
10 definitions
50
versionable survey benchmark definitions
Performance varies by domain and cohort calibration

Limitations

We believe in stating limitations honestly. Here's what doesn't work well yet.

Effect size magnitude is weak

HIGH

Mimiq often predicts the winning direction, but not by how much. Predicted effect sizes showed weak correlation with actual magnitudes (r = 0.04 on the curated 7-test suite), and only about 17% passed the broader suite’s realism threshold. Useful for ranking variants, not reliable for precise lift forecasting.

Selected cases are not an accuracy estimate

HIGH

The seven displayed cases all match direction, but they are a curated explanatory set. The broader 23-test fast run reached about 78% directional accuracy. A locked holdout and reproducible run manifest are required before making a generalized accuracy claim.

Domain coverage is uneven

MEDIUM

Performance is strongest on e-commerce, SaaS, and content engagement. Specialized B2B, medical, and government domains have higher variance in predictions.

Not using a custom model yet

MEDIUM

All results use a general-purpose LLM with Mimiq's proprietary behavioral system. A dedicated behavioral prediction model is in training and expected to improve both accuracy and calibration.

Headline A/B tests are hard

MEDIUM

Upworthy-style headline tests show the largest direction and effect-size errors. The model systematically under-predicts the magnitude of headline effects and sometimes picks the wrong winner.

Some responses fail parsing

LOW

A small percentage (typically 2-5%) of model responses cannot be parsed into valid actions. These are excluded from results, which adds noise and reduces effective sample size.

Custom model in development

The results above are the current behavioral-system baseline. We are developing a dedicated propensity model intended to improve accuracy and calibration; no improvement is assumed until it passes a locked holdout.

Anti-sycophancy training

The model learns to say "no" from real negative feedback: product reviews, complaints, and sycophancy research datasets.

Real outcome data

Ground truth from published RCTs and the benchmark suite documented on this page.

Rejection patterns

Teaching the model to produce realistic objections, not generic praise. The negative signal is the product.

References

1.Hillstrom, K. (2008). MineThatData E-Commerce Research Dataset.
2.Matias, J. N., & Munger, K. (2021). The Upworthy Research Archive. Nature Scientific Data.
3.Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157).
4.Ariely, D. (2008). Predictably Irrational. HarperCollins.
5.UK Behavioural Insights Team (2013). Applying behavioural insights to organ donation.
6.Hallsworth, M. et al. (2017). The Behavioralist As Tax Collector. NBER Working Paper 20007.
7.Linek, M. & Traxler, C. (2021). Framing Social Information. J. Public Economics, 200.
8.Unbounce (2024). Conversion Benchmark Report. 41,000 pages, 57M conversions.
9.Contentsquare (2025). Digital Experience Benchmark. 90B sessions across 3,800+ brands.
10.Argyle, L. P. et al. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis, 31(3).
11.Santurkar, S. et al. (2023). Whose Opinions Do Language Models Reflect? ICML 2023.
12.Pew Research Center (2024). Americans' Views on Technology & AI.
13.Gallup (2024). State of the Global Workplace Report.
14.Edelman (2024). Trust Barometer Report.
15.Hofstede, G. (2001). Culture's Consequences. Sage Publications.

Run your own benchmark.

Test your landing page, pricing, or campaign against a simulated audience.

Try Mimiq Free