Synthetic Respondents

Synthetic respondents are only as good as the humans grounding them.

User Intuition finds the people, interviews them in depth, and keeps the corpus calibrated as populations move. What you build on top of it is yours.

  • Depth interviews, not survey batteries
  • 4M+ panel across 50+ languages
  • 30/60/90-day recontact
The Problem

A simulated population inherits the thinness of its grounding.

Simulating a population is now straightforward. Write a persona, prompt a frontier model, generate as many respondents as the budget allows. The transcripts read plausibly, the themes look right, and the cost per respondent rounds to nothing. The difficulty is not producing synthetic respondents. It is knowing whether the ones you produced behave like the people they claim to represent.

They usually do not, and the failure is specific rather than general. We measured it directly. In The Synthetic Mirage in Market Research we ran 117 real voice interviews against 90 LLM-generated participants on the same protocol across three frontier models. The synthetic participants reproduced the language of qualitative research without reproducing its variance. Every real cohort contains people who disengage, refuse the premise, contradict themselves, or volunteer a granular complaint about a specific product. In our comparison, 55% of real interviews fell in the low transcript-quality band against 0% of synthetic, and 26% of real participants produced a refusal pattern against 0% of synthetic.

What collapses is the distribution, not the average. A synthetic population converges on the modal answer and deletes the tail, and the tail is where strategy lives. Seed the next generation on that output and the convergence compounds, because the model is now imitating its own priors rather than a population.

None of this is an argument against synthetic respondents. It is an argument about what they are conditioned on. Every synthetic participant in our comparison was persona-prompted, which independent work has since shown to be close to the weakest grounding available.

The Evidence

How much does grounding actually change accuracy?

Enough that it dominates every other design choice.

The measurement comes from outside our company. A Stanford-led team interviewed 1,052 people for two hours each using an AI interviewer, built one language-model agent per participant, then tested how accurately each agent reproduced its own source person across 150 General Social Survey items, the Big Five inventory, and five behavioral economic games. Because each agent is matched to a specific person, accuracy is normalized against how consistently that person reproduced their own answers two weeks later. A normalized accuracy of 1.00 means the agent predicts you as well as you predict yourself.

Agent grounded on Normalized accuracy
Depth interview and survey combined0.86
A depth interview transcript0.83
The person's own survey battery0.82
Demographic attributes0.74
A written persona paragraph0.71

Park et al., LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv:2411.10109 v3 (2026). Preprint. General Social Survey task, normalized accuracy.

The gap that matters commercially is the bottom of that table. A written persona paragraph — the method behind essentially every synthetic respondent sold today — reaches 0.71. Demographic prompting reaches 0.74. Anything grounded in what a person actually said about themselves lands twelve points higher. That distance is the difference between imitating a category and reproducing an individual, and it does not close with a better model or a larger sample.

The top of the table is worth reading precisely, because it is easy to overstate. Interview-grounded and survey-grounded agents did not differ from each other on this task; the authors report no significant difference between them. What did win was combining the two, at 0.86. So the honest claim is not that interviews beat surveys — it is that rich self-report grounding of either kind beats prompting, and that the two sources together beat either alone. Where interviews do separate is personality: on the Big Five inventory, interview-grounded agents reached 0.80 against 0.65 for survey-only, and adding survey data on top produced no further gain.

Interview length has sharply diminishing returns. Removing 80% of each transcript — 96 of the 120 minutes — still produced agents at 0.79. The remaining 24 minutes carries most of the signal, which means grounding does not require the two-hour research protocol that would make it uneconomic. It also means length should follow the research objective rather than a belief that longer is proportionally better.

Grounding also changes who the population misrepresents. Demographic prompting asks a model to infer an individual from a category, which is the mechanism that produces stereotyping. Measured as Demographic Parity Difference, agents grounded in self-report data of either kind were consistently less biased than demographic-grounded ones: political-ideology disparity on the General Social Survey fell from 13.75% to 8.60% for interview-grounded agents and 7.09% for combined agents. On the Big Five, interview-grounded agents cut disparity from 0.166 to 0.063 and combined agents to 0.048. Giving the model a person's own account removes the inference that generates the bias.

The Pipeline

Four stages, and we run the expensive ones.

The agent architecture is public. The grounding corpus is what costs money and time to produce.

01

Source the population

Stratified recruitment from a 4M+ opted-in panel across consumer and professional audiences in 50+ languages, with multi-layer verification on every participant. Stratification is set to the population being simulated rather than to a convenience sample, because a grounded corpus inherits whatever skew its recruitment carries.

02

Interview in depth

AI-moderated voice interviews run against a shared protocol, with adaptive follow-ups that probe rather than advance when an answer is thin. Hundreds to thousands run in parallel with consistent probing depth and uniform downstream coding — the constraint that historically capped depth interviewing at six to ten per researcher per week.

03

Structure the corpus

Transcripts ship alongside a latent-variable ontology: motivations, emotional drivers, identity claims, anticipated regret, perceived tradeoffs. Voice-derived features such as latency, hesitation, and certainty are extracted at processing time and surfaced as fields. Provenance travels at record level — study ID, protocol version, schema version, consent scope.

04

Recalibrate on a cadence

30, 60, and 90-day recontact against the same consented cohort. This is what a per-project panel cannot supply: populations drift, and the standard accuracy metric normalizes against each participant's own consistency over time — a denominator that cannot be computed without returning to the same people.

Who It's For

Three ways teams use a grounding corpus.

01

Independent validation

A held-out real cohort to measure a simulated population against. The value of this one depends on it being independent: a vendor validating its own simulation against ground truth it also sourced is grading its own work. For teams presenting simulation output to a fiduciary audience, third-party validation data is the part that survives scrutiny.

02

Recalibration supply

Ongoing depth interviews against a stable, re-contactable cohort for teams already running a simulation and recalibrating on a cadence. Fresh observations from the same people, not a fresh sample of different people — the distinction that makes drift measurable rather than merely suspected.

03

Corpus construction

The initial grounded population for teams building simulation capability rather than buying it. Commissioned to your stratification and protocol, delivered as transcripts plus structured ontology, with the recontact relationship preserved so the corpus can be maintained rather than rebuilt.

Procurement

What you license, how it's licensed.

Corpora are commissioned against your stratification and protocol, or licensed from existing studies where the domain already overlaps. Terms are the same as for our other data products.

Data formats
JSONL, Parquet, transcript bundles. Other formats on request.
Allowed uses
Fine-tuning, evaluation, reward modeling, synthetic data generation, retrieval-augmented generation, internal benchmarking, derivative datasets, model distillation. Use is bounded by the consent block on each record.
Licensing models
Exclusive, semi-exclusive, or syndicated. Pricing scales with corpus size, exclusivity, and grounding depth (number and type of behavioral mechanisms applied).
Retention & provenance
User Intuition retains ownership of the underlying corpus. Licensees receive structured derivative data under the terms of the agreement. Provenance is carried at the record level: study ID, protocol version, ontology schema version, consent scope.
Audio & video
Transcripts ship; audio and video do not. Voice-derived features are extracted at processing time and surfaced as ontology fields on the transcript. Buyers requiring biometric audio for specific use cases can discuss separately under additional consent terms.
Sample data
Anonymized illustrative samples available on inquiry.
Team

Who's running this.

User Intuition runs a self-serve research SaaS used by product, customer-experience, and insights teams to commission AI-moderated voice interviews and the structured outputs derived from them. The same platform infrastructure — including the Human Signal MCP, our model-context protocol layer for grounding agents in real human cognition — produces the human preference datasets described here.

The team is led by the founder and CEO (Harvard MBA, BS Electrical Engineering, Yale), and the broader User Intuition Research Team — the methodologists, engineers, and panel operators who design protocols, build the ontology, and run the studies.

Methodology, infrastructure, and team have produced tens of thousands of structured human-cognition interviews across consumer and professional domains since the platform was built.

FAQ

Grounding and validation questions.

Synthetic respondents are language-model agents that answer research questions in place of people. Their accuracy depends almost entirely on what they are grounded in. Park et al. (arXiv:2411.10109 v3, 2026) built one agent per participant from a national sample of 1,052 Americans and measured each against its own source person. Agents grounded on a written persona paragraph reached 0.71 normalized accuracy and demographic attributes 0.74, while agents grounded in that person's own self-report data reached 0.83 from a depth interview, 0.82 from a survey battery, and 0.86 from both combined. Normalized accuracy of 1.00 would mean the agent predicts the person as well as they predict themselves two weeks later. The grounding corpus, not the model, is the binding constraint.

Hold out a real cohort and compare distributions rather than averages. A simulated population never compared against real humans carries no error bar. Test the properties synthetic respondents lose first: the share of low-engagement responses, the rate of refusal or thesis rejection, the presence of outliers, and the frequency of specific named detail. In User Intuition's 117-real-versus-90-synthetic comparison, 55% of real interviews fell in the low transcript-quality band against 0% of synthetic, and 26% of real participants produced a refusal pattern against 0% of synthetic. Matching thematic content is not validation, because thematic content is the one thing synthetic respondents reliably reproduce.

Shorter than most teams assume. Park et al. (arXiv:2411.10109 v3, 2026) built agents from two-hour interviews, then removed 80% of each transcript — 96 of the 120 minutes — and the resulting agents still scored 0.79 normalized accuracy against their own source participants, close to the 0.83 achieved on the full transcript. Most of the predictive signal sits in the first stretch of an adaptive interview, so a grounding corpus does not require the two-hour research protocol that would make it uneconomic. Interview length should follow the research objective rather than an assumption that longer is proportionally better.

Because demographic prompting asks a model to infer an individual from a category, which is the mechanism that produces stereotyping. Park et al. (arXiv:2411.10109 v3, 2026) measured Demographic Parity Difference across subgroups and found agents grounded in self-report data consistently less biased than demographic-grounded ones. On the General Social Survey, political-ideology disparity fell from 13.75% for demographic agents to 8.60% for interview-grounded agents and 7.09% for combined agents. On the Big Five inventory it fell from 0.166 to 0.063 for interview-grounded agents and 0.048 for combined. Giving the model the person's own account replaces the inference that generates the bias.

For two reasons. Populations drift, so a corpus grounded once describes a moment rather than a market, and any team recalibrating on a cadence needs fresh observations from the same cohort. The second reason is structural: the standard accuracy metric normalizes agent performance against how consistently each participant reproduces their own answers over time. That denominator cannot be computed without going back to the same people. A panel sourced per-project on a cost-per-interview basis cannot supply it. User Intuition runs 30, 60, and 90-day recontact against a consented, quality-verified panel.

Real interviews, and the infrastructure that produces them at population scale. User Intuition sources participants, runs AI-moderated depth interviews in 50+ languages, and licenses the resulting transcripts and structured ontology as a grounding corpus. What a licensee builds on top of that corpus — synthetic respondents, calibration sets, evaluation holdouts — is their decision. The company's position, supported by its own published comparison, is that ungrounded synthetic respondents collapse population variance, and that the human interview sits upstream of any simulation rather than being replaced by it.