Synthetic Respondents vs Real Participants: A Decision Guide for Insights & Research Teams

When synthetic respondents sharpen research questions and when real participants are required, with a decision framework and benchmarks against human panels.

Maya Kennedy

Insights, research, product and strategy teams are being asked to test more concepts in more markets on a flat budget, and synthetic respondents promise answers in minutes. They can sharpen the questions a team asks, but any decision that carries money, health or reputation should rest on research with real people.

Teams reach for synthetic respondents largely because research with people has meant waiting weeks. There is another way to hear from real people quickly: AI-moderated interviews. An AI-moderated interview is a qualitative research method in which an AI moderator holds a voice-to-voice conversation with a participant, follows the researcher’s discussion guide and asks follow-up questions based on what that person says.

What are synthetic respondents?

Synthetic respondents are simulated research participants generated by a large language model (an AI system trained on text) and prompted to answer as a defined type of person. They also go by synthetic personas, AI respondents or synthetic consumers.

The clearest difference between synthetic respondents and research with real people is who gives the answers. When a streaming service tests a new price tier through AI-moderated interviews, the AI moderator only asks the questions and follow-ups, and a subscriber gives every answer. In a synthetic panel, no subscriber takes part, and a model writes what one might say.

AAPOR’s May 2026 report describes synthetic output as a model-based approximation of what a person might say, rather than a direct observation of human expression. ESOMAR, AAPOR, MRS and the Insights Association all treat synthetic respondents as something to use alongside research with real people, never in place of it.

How do the related terms differ?

The terms get used loosely, so these working definitions draw on ESOMAR and AAPOR:

  • Synthetic response: A synthetic response is a model-generated answer to a survey or interview question.
  • Synthetic people (synthetic personas): A synthetic persona is the simulated respondent itself, defined by a prompt or profile.
  • Synthetic research: Synthetic research is any study in which synthetic responses play a significant role in sampling, analysis or interpretation, and the ICC/ESOMAR Code requires disclosing it.
  • Synthetic data: ESOMAR defines synthetic data as “information that has been generated to replicate the characteristics of real-world data.”
  • Digital twin: A digital twin is a model grounded in one specific person’s prior surveys, interviews or behavior.
  • Synthetic sample (synthetic boost): A synthetic sample adds model-generated respondents to a human sample to fill out small segments. Boost only small segments, and only on top of an adequate human sample.
  • AI consumer panel: An AI consumer panel is a supplier’s standing library of synthetic personas, queried repeatedly instead of recruiting a fresh sample.

How are synthetic respondents built?

Synthetic respondents are built in three main ways, and the difference comes down to how much the model knows about an actual person before it answers. All three sit under the same label, but they do not perform alike:

  • A persona written from a profile: A researcher describes a type of person, such as age, attitudes and habits. The model then answers in character.
  • A persona given a person’s past answers, reviews or interview transcript: The model reads what that person actually said or did before it answers.
  • A persona built from one person’s full history of answers (a digital twin): Each person gets one simulated respondent. It is built from hundreds of that person’s own answers.

What does synthetic output miss?

However a persona is built, it returns one of three kinds of output, and all three share one flaw: they read fluently but miss the specific detail a person gives.

  • Simulated survey answers: Each persona returns a completed questionnaire. A travel app team can get back 500 completed questionnaires overnight, none of them from a traveler.
  • Synthetic open-ends and interviews: A persona writes a paragraph-length answer to an open question. Asked why it cancelled a streaming subscription, a persona writes a tidy paragraph about price and content. It cannot describe the show that left the catalog or the price email that made a particular subscriber give up, which is the detail an interview with a subscriber surfaces.
  • Persona-driven concept reactions: A persona rates a brand or concept on a scale. The ratings look plausible, but they show what a model expects a shopper to say, not what a shopper thinks.

Where do synthetic respondents help?

That missing detail matters least while a team is still deciding what to ask, which is where synthetic respondents earn their place. AAPOR’s task force flags serious validity risks for any use beyond “clearly labeled pretesting, pilot work, or exploratory diagnostics.” Four uses fit that boundary:

  • Early hypothesis generation: Personas can rank which hypotheses are worth testing with real people, leaving the size of any effect to the human study.
  • Exploring a space before the study: Personas can map the language and questions around a category before anyone writes the guide.
  • Message and concept pre-testing: Personas can cull broken variants so the human study takes fewer, stronger options to real people. A streaming service drafting messages for a new price tier, for example, can use personas to drop the unclear or contradictory ones and take only the clearest to subscribers. Treat direction of preference as the only readable output.
  • Stress-testing a discussion guide: Running the guide through personas catches leading questions, wrong order and jargon. Flag any question every persona answers identically, and any wording that hints at the answer the team hopes for.

If the output changes what a team asks, synthetic is fine as long as it is labeled. If it changes what the team decides, recruit real people.

Where do synthetic respondents mislead?

Once the output starts to shape what a team decides rather than what it asks, the limits show. Synthetic respondents fail where the model has no past to draw on, or where the answer depends on experience a text model lacks:

  • New or novel behaviors: When a brand enters a new category or launches a new product, the model has no past to draw on. In an HBS working paper, the model overestimated what people would pay for a projector more than threefold before it was trained on human answers.
  • Emotion and context: A model misses how a change in wording shifts how people feel about a message. Reword a scenario and people’s reactions move; a model’s barely do.
  • Sensory testing: A model has never tasted, smelled or handled anything. A snack brand can describe a new flavor to a model, but the rating that comes back reflects the description, not the taste.
  • Niche or hard-to-reach segments: The segments hardest to recruit are the ones models reflect worst. A model has seen little text from small or specialized groups, so it fills the gap with stereotype.
  • Anything a decision rests on: Relationships between variables do not hold up, and results shift even when the same prompt runs months apart. Bisbee et al. found that many relationships diverged from the human data (see the benchmarks below), and the same prompt run in April and July 2023 produced materially different results.
  • Model bias and invented findings: Models tend to favor the last answer option in a list and can report effects that do not exist in human data.

If the answer depends on new behavior, feeling, taste, group decisions or a thin segment, recruit real people.

How accurate are synthetic respondents?

Suppliers often answer these limits with an accuracy figure. Before trusting one, ask what was tested, on which topic and against which human sample, and ask for a test against people in the team’s own category and method. The two peer-reviewed comparisons below show why that matters:

Study (source, year) Human sample What was compared Result
Bisbee et al., Political Analysis, 2024 About 7,500 US election survey respondents Relationships between variables, synthetic vs human Nearly half significantly different; a third of those pointed the opposite way
Toubia et al., Marketing Science, 2025 About 2,000 US participants, 500+ questions Digital twin vs held-out answers About 72% accuracy, or about 88% of how consistently people matched their own earlier answers

Totals on well-documented topics can match closely; individual answers and the spread of opinion do not. The AAPOR task force says the field “has not yet reached consensus on when, if ever, AI-generated responses can stand in for human ones.” If an accuracy claim has no test against people in the team’s own category behind it, treat it as unproven.

Does synthetic research save time and money?

Even where accuracy is unproven, the case for synthetic respondents usually rests on speed and cost. Synthetic output does buy speed, but the savings are unproven and the human study still needs a plan:

  • Speed: Synthetic output comes back in minutes or hours instead of weeks, though a fast answer that points the wrong way costs more than a slow one.
  • Cost: No independent, controlled study quantifies savings against a matched human panel; the percentage claims in circulation come from suppliers.
  • Human data: Every grounded method in the glossary and construction list needs human data first.

A travel app team screening five new feature concepts can get a synthetic screen back in an afternoon that narrows the shortlist. The time saved sits on the shortlist, not on the decision, so the team still tests the surviving concepts with real people before committing to a launch. That human study is the step teams expect to take weeks, and with AI-moderated interviews it no longer has to; the last section explains how.

If synthetic output saves time, count it on the shortlist. If the result sets the launch, the time goes to people.

Which studies should stay with real people?

The failure modes above map onto specific research methods. For the methods below, the evidence points one way: keep them with real people.

Methodology Fit for synthetic respondents Why
Conjoint, discrete choice, willingness to pay Keep human In the HBS study, the model valued removing aluminum from deodorant at about +$0.90, while people in an earlier human study valued it at about -$2: the model got the direction wrong.
MaxDiff Keep human No peer-reviewed study has validated synthetic MaxDiff.
Monadic concept testing, purchase intent Human for the decision; synthetic only to cull Relationships between answers diverge from human data (see the benchmarks).
Sensory evaluation Never A model can read a product description but cannot taste, smell or handle the product.
Zero-to-one innovation Keep human With no past behavior to learn from, the model guesses (see the projector example above).
Low-incidence B2B Keep human No peer-reviewed study validates synthetic IT buyers, procurement leads or C-suite.

How do you check synthetic output?

Where synthetic output does fit, mostly in shaping questions and culling options before the human study, it still needs checking. Run these checks before any synthetic output leaves the research team:

  1. Benchmark against people on the team’s own study. Run a small human study alongside and measure synthetic accuracy against how often people repeat their own earlier answers when asked again, the ratio Toubia et al. use.
  2. Compare the full spread of answers as well as the averages. Averages can look right while subgroups, and even the direction of an effect, are wrong.
  3. Report worst-subgroup results. Pick the subgroups that matter before the run, and report the worst one, not the average.
  4. Reverse the answer order and rerun. If answers flip with the order, the output is reading the list, not the question.
  5. Rerun weeks apart. Run the same prompt again after a few weeks; if results shift materially, as they did for Bisbee et al., treat the output as unstable.
  6. Correct against a human sample. A smaller human sample can measure how far the synthetic output is off and correct for it. Treat the correction as a check, not a substitute for the human study.

How can a team test whether a study needs real participants?

A streaming service weighing a new price for its annual plan wants to know whether subscribers will pay more, and synthetic personas are cheap and fast. The earlier sections covered where synthetic output helps and where it fails by method. This section turns that into a quick test any planned study can be run through, using five questions:

  1. What changes because of the result? A question, guide or shortlist can rest on labeled synthetic output; a price, claim, launch, formulation or CMO slide needs real participants.
  2. Has the model seen this behavior before? New categories, habits and competitive contexts sit outside the training data, so the model guesses.
  3. Does the answer depend on feeling, tasting or watching other people? Emotion and sensory experience fail in the benchmarks, and no published study yet tests group decisions such as a buying committee.
  4. Is the segment thin or specialized? Low-incidence and professional groups collapse toward stereotype, so synthetic boosting is defensible only on top of an adequate human base (see the glossary).
  5. Will anyone outside the team see it? Disclosure obligations apply, and the reputational cost of an undisclosed synthetic result falls on the team.

[VISUAL: The five-question decision framework as a one-page flow, with each question as a yes or no step leading to either ‘labeled synthetic is fine’ or ‘recruit real participants’]

When the five questions point to real people, the study no longer has to take weeks. Strella runs qualitative, AI-moderated interviews in which an AI moderator holds a voice-to-voice conversation with each participant and asks follow-up questions based on what that person said. Screener answers are fact-checked, and every insight links back to the interview and quote it came from. Teams can choose AI, human or hybrid moderation, and the researcher still owns the question, the guide and the interpretation.

Bring your research question to us and we will talk it through. See Strella in action

FAQ

What are synthetic respondents in market research?

Synthetic respondents are AI-generated stand-ins for research participants, prompted to answer as a defined persona. The ICC/ESOMAR Code requires researchers to disclose their use in a study.

Are synthetic respondents accurate enough for concept testing?

Use real participants for a go or no-go decision. Published comparisons find that individual answers and the spread of opinion diverge from what people say. Use them to screen out weak concepts, then test the survivors with people.

When should I use real participants instead of synthetic respondents?

Whenever the result will set a price or support a claim, formulation, launch or board recommendation. Novel behavior, emotion, taste or smell, group interaction and niche segments also call for real participants.

What are the bias and legal risks of synthetic respondents?

Researchers may see last-option bias, socially desirable answers, stereotyped minority groups and effects that do not appear in human data. Legally, a privacy team should confirm that panel or customer data can be used to train or ground a model.

Can synthetic and human research work together?

Yes: synthetic first to shape hypotheses and test the guide, people second for the study itself. Disclose any blend and check any synthetic boost against a human holdout.

Do synthetic respondents mean I need fewer researchers?

  1. Synthetic output needs more researcher judgment, not less: someone has to choose the human benchmark, read the worst-subgroup results and decide what the output is allowed to change.