How Accurate Are Synthetic Users? A Preregistered Study
A floor-to-ceiling test of what conditioning data an LLM synthetic respondent needs — hypotheses frozen before any data was generated.
How accurate are synthetic users? In the first preregistered study of its kind, User Intuition measured what makes LLM synthetic respondents faithful to the real people they simulate. Using 53 consumer panelists, 265 held-out interview questions, and 1,625 scored predictions evaluated by two independent blind judges, we compared conditioning methods on a floor-to-ceiling scale — from a random stranger to the person's own full interview. Demographic-only personas scored just 19% of the way to the ceiling: barely better than a stranger. Conditioning on the person's real interview transcript closed 81% of the fidelity gap (95% CI, 68–96%). An LLM-generated summary of the interview retained 91% of full-transcript fidelity. Accuracy was weakest on reasoning — why people do what they do — exactly the dimension research exists to uncover. The conclusion: interviews, not demographics, are what make synthetic respondents accurate.
Executive Summary
"How accurate are synthetic users?" is the question every insights team asks before buying one, and the honest answer until now has been that nobody knows. Published accuracy claims range from roughly 37% to 95%, and none of them are comparable: different tasks, no floor to measure against, no ceiling to measure toward, and — critically — no isolation of the one variable that plausibly matters, which is what the synthetic respondent was conditioned on. So we ran the study. We preregistered our hypotheses and froze them before a single prediction was generated, then used our own consumer panel as the ground truth: 53 members who had each completed a recorded onboarding interview, 265 held-out question–answer pairs stratified across five laddering rungs, and 1,625 predictions scored blind by two frontier-LLM judges from different vendors. Every result is reported on a floor-to-ceiling scale — 0% is a different member's synthetic answering the same question (a stranger), 100% is a synthetic conditioned on the person's own full interview transcript. The finding that matters is not a single accuracy number. It is that the interview, not the demographic profile, is the active ingredient: demographics alone recover 19% of the distance from stranger to ceiling, while the person's real interview content closes 81% of the remaining fidelity gap. One preregistered hypothesis found no support, and we publish it here, because that is what preregistration is for.
- Demographic-only synthetic respondents barely beat a stranger: they land 19% of the way from a random-stranger floor to the person's own full-interview ceiling.
- The real interview is the active ingredient: full-transcript conditioning closes 81% of the fidelity gap (95% CI, 68–96%), replicating Park et al. (2024) on open-ended interview data.
- Every layer of interview content helps, at every question depth: behavior-only excerpts reach 58% of the gap, adding attribute and functional content reaches 81%, and the full transcript defines the 100% ceiling.
- Summaries are a cheap, close substitute: an LLM-generated summary of the interview retains 91% of full-transcript fidelity, with no statistically detectable concentration of loss on emotional or value questions.
- A preregistered hypothesis came back inconclusive and we publish it: we predicted interview depth would matter specifically for deep questions, and the measured interaction was −0.17 (95% CI, −0.50 to 0.15), spanning zero.
- Fidelity is weakest exactly where research matters most: the reasoning dimension — why people do what they do — scores lowest in every conditioning arm, and demographics-only reasoning sits barely above the stranger floor.
Why No Two Synthetic-User Accuracy Numbers Mean the Same Thing
Published accuracy claims for synthetic respondents span roughly 37% to 95%. The spread is not a disagreement about the answer — it is the absence of a shared question.
Three things missing from every public number
The first missing piece is a floor. An accuracy score of 70% sounds impressive until you learn what a completely uninformed guess would have scored on the same task. Many consumer questions have a dominant answer; a synthetic respondent that knows nothing about the person can land a large share of them by regression to the population mean. Without a floor, the score measures the predictability of the question set, not the fidelity of the simulation.
The second missing piece is a ceiling. Even a perfect model of a person cannot predict what that person will say with 100% reliability, because people are not internally consistent from one moment to the next. A fidelity number is only interpretable relative to the best achievable performance on that specific task — which means you need an upper anchor built from the person's own data.
The third missing piece is isolation. Vendors describe their synthetic respondents as "grounded" or "data-backed" without specifying what data, in what quantity, contributes what. If a persona is built from demographics plus purchase history plus a survey plus an interview, and it scores well, the score tells a buyer nothing about which ingredient carried the weight — or whether the cheap ingredients could be dropped.
What we did instead
We built a floor and a ceiling from the same people, held out real answers those people had already given, and varied exactly one thing: what the synthetic respondent was allowed to know. Every number on this page is a position on that floor-to-ceiling scale, not a naked accuracy percentage. The hypotheses were preregistered and frozen before any prediction was generated, with an amendment log kept inside the preregistration document.
This is wave 1 of an ongoing benchmark. It answers a narrower question than "are synthetic users accurate" — it answers "how much of a real person's held-out interview answers can a synthetic respondent reproduce, and which conditioning data gets it there."
- Public synthetic-respondent accuracy claims range from roughly 37% to 95% with no shared floor, ceiling, or isolation of the conditioning variable — which makes them non-comparable rather than contradictory.
How We Measured Fidelity
A floor-to-ceiling design on 53 panel members, 265 held-out answers, seven conditioning arms, and two blind judges cross-checked against blinded human scoring.
The sample
Fifty-three US and Southeast Asian consumer panel members, each with a recorded onboarding interview (voice, AI-moderated) and self-reported signup demographics. Eleven further members were excluded before any analysis ran: 6 with no attributable interview, 2 immediate hang-ups, and 3 with content-contradicted records (identity-level mismatches between the interview and the profile). Those exclusion rules were written into the preregistration before generation began. Member-to-interview attribution had to be reconstructed — the source table carried a stamping bug — via timestamp and email matching, cross-checked content-blind; the validation statistics are recorded in the preregistration.
The held-out design
For each member we held out 5 real exchanges from their interview, stratified across the five laddering rungs of a depth ladder: behavior, attribute, functional, emotional, and value. That yields 265 held-out questions. Rung tagging was done by an LLM tagger and checked against two human raters — the human–human ceiling on this task was 80%, the tagger agreed with the primary rater 84% of the time, and 100% on items where the two humans agreed. Held-out exchanges and their conversational neighbors were removed from every conditioning input, so no arm could see the answer it was being asked to predict. All arms answered the same questions.
The seven arms
The floor (A0) is a different member's fully-conditioned synthetic answering the same question — a well-briefed stranger. The ceiling (A5) is a synthetic conditioned on the member's own full interview transcript. Between them: demographics only (A1); demographics plus professional history (A2, exploratory); behavior-only interview excerpts (A3); behavior plus attribute and functional content (A4); and an LLM-generated summary of the interview (A6).
Scoring and judge agreement
Two frontier-LLM judges from different vendors, blind to which arm produced each answer, scored every prediction on three dimensions — claims, stance, and reasoning — at 0–2 each. Judge output was cross-checked against 50 blinded human-scored pairs: within-one-point agreement was 100% for judge A and 99% for judge B; exact agreement was 73% for both. Mean inter-judge disagreement was 0.61 summed across the three 0–2 dimensions. Disagreements are absorbed by averaging the two judges, with no third-pass adjudication, and all agreement statistics are disclosed rather than summarized.
Inference discipline
Confidence intervals come from a bootstrap that cluster-resamples members, 2,000 draws. Only three preregistered pooled contrasts are treated as inferential; every cell-level value carries its own n and is explicitly descriptive. Cells below an n=10 display floor are suppressed. Two robustness checks were run: dropping the 36 held-out questions that were ASR fragments left the arm ordering unchanged, and regenerating 60 pairs with a different generation model moved scores by no more than 0.25 raw points on the 0–6 scale — the results are not an artifact of the model doing the generating.
Scope
All claims here are scoped to within-domain response prediction: reproducing answers a person gave inside an interview whose other parts the model can see. This is wave 1 of an ongoing benchmark; future waves are preregistered before they run, and each wave publishes its nulls alongside its positive results.
- Two frontier-LLM judges from different vendors, blind to arm, were cross-checked against blinded human scoring: within-one-point agreement 100% and 99%, exact agreement 73% and 73%.
- Rung tagging is noisy by nature: the human–human ceiling on the tagging task was 80%, and the LLM tagger reached 84% against the primary rater and 100% on human-consensus items.
- Results survive two robustness checks: dropping 36 ASR-fragment questions left arm ordering unchanged, and a 60-pair generation-model rerun moved scores by ≤0.25 raw points on the 0–6 scale.
Results: The Conditioning Ladder
Each rung adds one layer of what the synthetic respondent knows about the person. The jump that matters is the first layer of real interview content.
How Much of the Fidelity Gap Each Conditioning Level Closes
Position on the floor-to-ceiling scale — 0% = a well-briefed stranger, 100% = the person's own full interview transcript
Bar labels show the number of panel members contributing to each cell (n). A0 and A5 are the scale anchors by construction. A2 — demographics plus professional history — is suppressed: n=7 members, below the n=10 display floor, reported qualitatively only. Cell values are descriptive; only the preregistered pooled contrasts carry confidence intervals.
Reading the table
Every value below is a position on the floor-to-ceiling scale: 0% is a well-briefed stranger answering the question, 100% is a synthetic conditioned on the person's own full interview transcript. Each arm was scored on the same 265 held-out questions across the same 53 members, so the rows are directly comparable to each other in a way that cross-vendor accuracy claims are not.
A0, the random-stranger floor, sits at 0% by construction — it is the anchor, not a result. A1, demographics only, reaches 19% (n=53 members). A3, behavior-only interview excerpts, reaches 58% (n=53). A4, behavior plus attribute and functional content, reaches 81% (n=53) — A4's 81% is its position on the ladder, numerically coincidental with the 81-percentage-point gap between demographics and the full transcript reported below. A5, the full transcript, defines the 100% ceiling (n=53). A6, an LLM-generated summary of the interview, reaches 91% (n=53) — nine points below the transcript it was compressed from. A2, the exploratory professional-history arm, is suppressed: n=7 members, below the n=10 display floor, reported qualitatively only.
The headline: 81% of the fidelity gap
The preregistered contrast is A5 minus A1 — what the person's real interview adds on top of their demographic profile. It closes 81% of the fidelity gap (95% CI, 68–96%). That is the number that answers "how accurate are synthetic users," and its shape matters more than its size: accuracy is not a property of the model, it is a property of what the model was given. The same generation setup produces a near-stranger or a near-replica depending entirely on whether it has seen the person talk. This replicates Park et al. (2024) — who found interview-conditioned agents substantially outperformed demographic conditioning — on open-ended interview answers rather than structured survey instruments.
Depth is monotonic, not thresholded
There is no cliff and no plateau. Adding interview content improves prediction at every step, on every question depth, including the emotional and value rungs where a skeptic would expect simulation to break down: on those rungs specifically the ladder runs 62% → 80% → 100%. Nothing in the data suggests you can stop halfway up.
Where fidelity is weakest
Broken out by scoring dimension, the pattern is uncomfortable. Claims — what the person says they do — are the easiest to reproduce. Stance follows. Reasoning, the dimension that captures why a person does what they do, scores lowest in every single arm, and in the demographics-only arm it sits barely above the stranger floor. That is precisely the layer of an interview that qualitative research exists to surface, and it is the layer a demographic persona is least able to invent.
The professional-history arm, qualitatively
The A2 arm — demographics plus professional history — cannot be reported numerically at n=7. The qualitative observation is retained because it is a design warning rather than a result: profile-primed personas frequently answered consumer questions in a professional register, drifting off-domain into the voice of a job title rather than a shopper.
- Full-transcript conditioning closes 81% of the fidelity gap over demographics alone (95% CI, 68–96%) — the preregistered headline contrast.
- The ladder is monotonic: 19% (demographics) → 58% (behavior-only excerpts) → 81% (+ attribute and functional content) → 100% (full transcript).
- An LLM summary of the interview reaches 91% of full-transcript fidelity — most of the value of the transcript at a fraction of the context.
- The professional-history arm is reported qualitatively only (n=7, below the n=10 display floor): profile-primed personas frequently answered consumer questions in a professional register.
The Null We're Publishing
We predicted that interview depth would matter more for deep questions. It came back inconclusive. Preregistration is what makes that sentence publishable rather than quietly deleted.
What we predicted
Hypothesis 1, frozen before any generation ran, was a differential-depth claim: that the improvement from adding deeper interview content would be steeper for emotional and value questions than for behavior questions. The intuition is familiar to anyone who has run interviews — you can guess what someone buys, but not why it matters to them, unless they told you.
What we found
The measured interaction was −0.17, with a 95% confidence interval of −0.50 to 0.15. The interval spans zero. The comparison cannot be determined either way from this sample, and we do not claim the reverse either — this is an inconclusive result, not evidence that depth matters less for deep questions.
Two disclosed caveats point the same direction. Rung-tag noise attenuates exactly this interaction: the tagger reached 84% against the primary rater on a task where the measured human–human ceiling is 80%, and misallocated rungs blur precisely the contrast H1 depends on. And n=53 is underpowered for interaction detection even under clean tagging. The honest reading is that this study cannot resolve the question, not that the question is settled.
What survived
The main-effect result is untouched by the null. More interview content improved prediction at every question depth, monotonically, including on emotional and value questions. What did not survive is the sharper, more sellable version of that claim — that deep questions specifically require deep interviews. We are not making that claim, on this page or anywhere else.
A second, smaller null belongs here too. Summaries retain 91% of full-transcript fidelity overall and 84% on emotional and value questions, which looks like the loss concentrating on deep questions. It is not statistically supported: the ratio difference is 0.073 with a 95% CI of −0.07 to 0.19. The defensible claim is "summaries retain roughly 91%, with no demonstrated concentration of loss."
Why this section exists
A preregistered study that only reports its confirmed hypotheses is an unregistered study with extra paperwork. The reason to freeze hypotheses before generating data is to make the inconvenient outcome publishable, and the only way that discipline means anything to a reader is if they can see it operate at least once. This is that once.
- H1 (depth-specificity) is inconclusive: measured interaction −0.17, 95% CI −0.50 to 0.15, spanning zero. Neither the predicted direction nor its reverse is supported.
- The apparent concentration of summary-conditioning loss on emotional and value questions is not statistically supported: ratio difference 0.073, 95% CI −0.07 to 0.19.
What This Means for Synthetic Research
Two things called "synthetic respondents" are not the same product. One is a demographic guess; the other is a compression of a real conversation.
Demographic personas are a different category, not a cheaper tier
A synthetic respondent conditioned only on age, gender, region, and income lands 19% of the way from a stranger to the real person. That is not a discount version of a real participant — it is a differently-shaped object that happens to produce fluent sentences in the same format. It reproduces the population mean with a name attached. Our earlier study, The Synthetic Mirage in Market Research, documented the failure modes qualitatively; this one puts a controlled number under them.
Interview-grounded synthetics are a real capability with a scoped use
Conditioned on the person's own interview, the same machinery closes 81% of the fidelity gap (95% CI, 68–96%). That is a useful object — for extending a completed study, pressure-testing a stimulus against people you already interviewed, or letting a stakeholder ask a follow-up question at 2am. It is not a replacement for talking to someone new, and this study cannot tell you how it behaves on a person in an unfamiliar context, because every held-out question came from inside an interview the model could partly see. See how we build interview-grounded synthetic respondents for where we do and do not use them.
Summaries make it affordable, and the affordability is real
A summary of the interview retains 91% of full-transcript fidelity. For anyone running this at scale, that is the difference between a workable context budget and an unworkable one — and the 9-point cost is measured rather than assumed. There is no demonstrated concentration of that loss on emotional or value questions.
The bridge to the academic work
Park et al. (2024) reached the same structural conclusion from a different direction: interview-conditioned agents outperformed demographic conditioning on structured instruments. The broader academic literature on simulation benchmarks and identity flattening has been converging on this for two years while commercial content has largely ignored it. Our contribution is narrow and specific — the same result on open-ended interview answers, with a floor and a ceiling, preregistered, with the null published. Related reading: agentic research versus synthetic panels and real versus synthetic research participants.
The practical test for a buyer
Ask any vendor quoting an accuracy number three questions. What is your floor — what would an uninformed guess have scored? What is your ceiling — what is the best score achievable on this task? And what exactly was the synthetic respondent conditioned on? If the answer to the third question is "demographics," the first two answers will not save it.
- Demographic personas and interview-grounded synthetic respondents are separated by 81% of the fidelity gap (95% CI, 68–96%) — a category difference, not a quality tier.
- Summary conditioning retains 91% of full-transcript fidelity, making interview-grounded synthetics practical within realistic context budgets.
Limitations, Honestly
Everything on this page is scoped to within-domain response prediction on a single panel of 53 people. Here is what that rules out.
The holdout is inside the interview
Held-out answers came from interviews whose remaining parts the model could see. That inflates absolute fidelity relative to predicting how a person would respond in a novel context — a new category, a new stimulus, a different mood six months later. This is the main reason relative lift, not absolute accuracy, is the headline metric: the inflation applies roughly equally across arms, so the comparison between arms survives what the absolute number does not.
n=53, one panel, heterogeneous topics
Fifty-three members from a single panel, with interviews spanning different topics. That is enough to detect a large main effect with a clustered bootstrap and nowhere near enough to detect an interaction — which is exactly what happened to H1. Treat every per-cell value as descriptive.
Panel demographic records carry real-world noise
Self-reported signup demographics include inflation and malformed fields. That noise attenuates the demographics arm, so 19% is arguably a slightly pessimistic reading of what a clean demographic profile could do. It is also the point: that same noise is a property of every demographic-persona product, which builds on the same kind of self-reported data. Nobody gets a clean version of this input in production.
Tag noise and the exclusion effect
Rung tagging runs at 84% against the primary rater on a task with a measured 80% human–human ceiling, which attenuates every rung-level contrast including H1. Separately, excluding 3 content-contradicted members biases the retained sample toward demographic consistency — which means the weak demographics result survives its most obvious threat: we removed the members most likely to make demographics look bad, and demographics still landed at 19%.
What this study does not measure
It does not measure whether synthetic respondents produce good research conclusions, only whether they reproduce individual held-out answers. It does not measure novel-context generalization. It does not measure any vendor's product other than the conditioning ladder described in the methods. And it is wave 1 — the numbers here are a first controlled reading, not a settled benchmark.
- All claims are scoped to within-domain response prediction: reproducing held-out answers from inside an interview the model can partly see.
- Excluding 3 content-contradicted members biases the retained sample toward demographic consistency — the weak demographics result survives its main threat.
Implications & Recommendations
The study answers a buying question, not just a scientific one. If you are evaluating synthetic respondents — or being sold them — these are the implications that follow directly from the measurements above.
- 1 Treat any accuracy number without a floor and a ceiling as uninterpretable. A raw accuracy percentage measures how predictable the question set is at least as much as how faithful the simulation is. Ask what an uninformed guess scores on the same task, and what the best achievable score is. Published claims spanning 37% to 95% are not disagreeing with each other — they are measuring different things.
- 2 Ask what the persona is conditioned on before you ask how accurate it is. Conditioning explains the outcome. Demographics alone land 19% of the way from a stranger to the real person; the person's own interview closes 81% of the fidelity gap (95% CI, 68–96%). Same models, same prompts, opposite results. The conditioning input is the product.
- 3 Use synthetic respondents to extend research you already ran, not to skip it. Fidelity was measured on held-out answers from inside an interview the model could partly see. That supports follow-up questions against people you have already interviewed. It does not support simulating a population you have never spoken to, and this study offers no evidence either way about novel-context performance.
- 4 Budget for summaries, not full transcripts. An LLM-generated summary retains 91% of full-transcript fidelity, with no demonstrated concentration of loss on emotional or value questions. For teams running this across hundreds of participants, the 9-point cost buys a context budget that works at scale — and it is measured, not assumed.
- 5 Discount any synthetic finding that turns on 'why.' Reasoning is the lowest-scoring dimension in every conditioning arm, and in the demographics-only arm it sits barely above the stranger floor. Motivation, trade-off logic, and the reasons behind a preference are the weakest part of the simulation — and the part most research is commissioned to explain.
Frequently Asked Questions
Ready to understand your customers this deeply?
Preview a real study output — or start free and launch in minutes.
You only pay for quality interviews.
Every interview is automatically scored against your brief. Misses aren't charged.
No contract · No retainers · First insights in 24 hours