Synthetic respondents are only as good as the humans grounding them.
User Intuition finds the people, interviews them in depth, and keeps the corpus calibrated as populations move. What you build on top of it is yours.
- Depth interviews, not survey batteries
- 4M+ panel across 50+ languages
- 30/60/90-day recontact
A simulated population inherits the thinness of its grounding.
Simulating a population is now straightforward. Write a persona, prompt a frontier model, generate as many respondents as the budget allows. The transcripts read plausibly, the themes look right, and the cost per respondent rounds to nothing. The difficulty is not producing synthetic respondents. It is knowing whether the ones you produced behave like the people they claim to represent.
They usually do not, and the failure is specific rather than general. We measured it directly. In The Synthetic Mirage in Market Research we ran 117 real voice interviews against 90 LLM-generated participants on the same protocol across three frontier models. The synthetic participants reproduced the language of qualitative research without reproducing its variance. Every real cohort contains people who disengage, refuse the premise, contradict themselves, or volunteer a granular complaint about a specific product. In our comparison, 55% of real interviews fell in the low transcript-quality band against 0% of synthetic, and 26% of real participants produced a refusal pattern against 0% of synthetic.
What collapses is the distribution, not the average. A synthetic population converges on the modal answer and deletes the tail, and the tail is where strategy lives. Seed the next generation on that output and the convergence compounds, because the model is now imitating its own priors rather than a population.
None of this is an argument against synthetic respondents. It is an argument about what they are conditioned on. Every synthetic participant in our comparison was persona-prompted, which independent work has since shown to be close to the weakest grounding available.
How much does grounding actually change accuracy?
Enough that it dominates every other design choice.
The measurement comes from outside our company. A Stanford-led team interviewed 1,052 people for two hours each using an AI interviewer, built one language-model agent per participant, then tested how accurately each agent reproduced its own source person across 150 General Social Survey items, the Big Five inventory, and five behavioral economic games. Because each agent is matched to a specific person, accuracy is normalized against how consistently that person reproduced their own answers two weeks later. A normalized accuracy of 1.00 means the agent predicts you as well as you predict yourself.
| Agent grounded on | Normalized accuracy |
|---|---|
| Depth interview and survey combined | 0.86 |
| A depth interview transcript | 0.83 |
| The person's own survey battery | 0.82 |
| Demographic attributes | 0.74 |
| A written persona paragraph | 0.71 |
Park et al., LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv:2411.10109 v3 (2026). Preprint. General Social Survey task, normalized accuracy.
The gap that matters commercially is the bottom of that table. A written persona paragraph — the method behind essentially every synthetic respondent sold today — reaches 0.71. Demographic prompting reaches 0.74. Anything grounded in what a person actually said about themselves lands twelve points higher. That distance is the difference between imitating a category and reproducing an individual, and it does not close with a better model or a larger sample.
The top of the table is worth reading precisely, because it is easy to overstate. Interview-grounded and survey-grounded agents did not differ from each other on this task; the authors report no significant difference between them. What did win was combining the two, at 0.86. So the honest claim is not that interviews beat surveys — it is that rich self-report grounding of either kind beats prompting, and that the two sources together beat either alone. Where interviews do separate is personality: on the Big Five inventory, interview-grounded agents reached 0.80 against 0.65 for survey-only, and adding survey data on top produced no further gain.
Interview length has sharply diminishing returns. Removing 80% of each transcript — 96 of the 120 minutes — still produced agents at 0.79. The remaining 24 minutes carries most of the signal, which means grounding does not require the two-hour research protocol that would make it uneconomic. It also means length should follow the research objective rather than a belief that longer is proportionally better.
Grounding also changes who the population misrepresents. Demographic prompting asks a model to infer an individual from a category, which is the mechanism that produces stereotyping. Measured as Demographic Parity Difference, agents grounded in self-report data of either kind were consistently less biased than demographic-grounded ones: political-ideology disparity on the General Social Survey fell from 13.75% to 8.60% for interview-grounded agents and 7.09% for combined agents. On the Big Five, interview-grounded agents cut disparity from 0.166 to 0.063 and combined agents to 0.048. Giving the model a person's own account removes the inference that generates the bias.
Four stages, and we run the expensive ones.
The agent architecture is public. The grounding corpus is what costs money and time to produce.
01
Source the population
Stratified recruitment from a 4M+ opted-in panel across consumer and professional audiences in 50+ languages, with multi-layer verification on every participant. Stratification is set to the population being simulated rather than to a convenience sample, because a grounded corpus inherits whatever skew its recruitment carries.
02
Interview in depth
AI-moderated voice interviews run against a shared protocol, with adaptive follow-ups that probe rather than advance when an answer is thin. Hundreds to thousands run in parallel with consistent probing depth and uniform downstream coding — the constraint that historically capped depth interviewing at six to ten per researcher per week.
03
Structure the corpus
Transcripts ship alongside a latent-variable ontology: motivations, emotional drivers, identity claims, anticipated regret, perceived tradeoffs. Voice-derived features such as latency, hesitation, and certainty are extracted at processing time and surfaced as fields. Provenance travels at record level — study ID, protocol version, schema version, consent scope.
04
Recalibrate on a cadence
30, 60, and 90-day recontact against the same consented cohort. This is what a per-project panel cannot supply: populations drift, and the standard accuracy metric normalizes against each participant's own consistency over time — a denominator that cannot be computed without returning to the same people.
Three ways teams use a grounding corpus.
01
Independent validation
A held-out real cohort to measure a simulated population against. The value of this one depends on it being independent: a vendor validating its own simulation against ground truth it also sourced is grading its own work. For teams presenting simulation output to a fiduciary audience, third-party validation data is the part that survives scrutiny.
02
Recalibration supply
Ongoing depth interviews against a stable, re-contactable cohort for teams already running a simulation and recalibrating on a cadence. Fresh observations from the same people, not a fresh sample of different people — the distinction that makes drift measurable rather than merely suspected.
03
Corpus construction
The initial grounded population for teams building simulation capability rather than buying it. Commissioned to your stratification and protocol, delivered as transcripts plus structured ontology, with the recontact relationship preserved so the corpus can be maintained rather than rebuilt.
What you license, how it's licensed.
Corpora are commissioned against your stratification and protocol, or licensed from existing studies where the domain already overlaps. Terms are the same as for our other data products.
- Data formats
- JSONL, Parquet, transcript bundles. Other formats on request.
- Allowed uses
- Fine-tuning, evaluation, reward modeling, synthetic data generation, retrieval-augmented generation, internal benchmarking, derivative datasets, model distillation. Use is bounded by the consent block on each record.
- Licensing models
- Exclusive, semi-exclusive, or syndicated. Pricing scales with corpus size, exclusivity, and grounding depth (number and type of behavioral mechanisms applied).
- Retention & provenance
- User Intuition retains ownership of the underlying corpus. Licensees receive structured derivative data under the terms of the agreement. Provenance is carried at the record level: study ID, protocol version, ontology schema version, consent scope.
- Audio & video
- Transcripts ship; audio and video do not. Voice-derived features are extracted at processing time and surfaced as ontology fields on the transcript. Buyers requiring biometric audio for specific use cases can discuss separately under additional consent terms.
- Sample data
- Anonymized illustrative samples available on inquiry.
Who's running this.
User Intuition runs a self-serve research SaaS used by product, customer-experience, and insights teams to commission AI-moderated voice interviews and the structured outputs derived from them. The same platform infrastructure — including the Human Signal MCP, our model-context protocol layer for grounding agents in real human cognition — produces the human preference datasets described here.
The team is led by the founder and CEO (Harvard MBA, BS Electrical Engineering, Yale), and the broader User Intuition Research Team — the methodologists, engineers, and panel operators who design protocols, build the ontology, and run the studies.
Methodology, infrastructure, and team have produced tens of thousands of structured human-cognition interviews across consumer and professional domains since the platform was built.