All posts

Synthetic users: the complete guide

Tania ClarkeAugust 17, 202614 min
Webb field of candidate young brown dwarfs outside the Milky Way.
Guide

Synthetic Users

no citation, no claim

Webb · NGC 602 · 2024 · NASA
On this page

TL;DR

A synthetic user is an AI-generated profile that stands in for a real research participant. You can ask it questions, run a concept past it, or use it to pressure-test a document before taking that work to real people.

The most important distinction is whether the synthetic user is grounded or ungrounded. A grounded synthetic user is built from research you already own, with claims tied back to real interviews, surveys, support conversations, or other customer evidence. An ungrounded one starts with a prompt and fills in the gaps from the model's training data.

The question everyone asks is how accurate synthetic users are. The problem is that there is no stable human baseline to measure them against. People do not give perfectly consistent answers themselves, so a single accuracy percentage can sound much more definitive than the evidence really is. In a Stanford-led study of 1,052 people, participants reproduced their own prior answers about 81% of the time, while synthetic replicas matched the original answers about 68% of the time. The often-quoted 85% figure comes from normalizing that 68% against the 81% human self-consistency rate.

The better question is what synthetic users are reliably useful for. The strongest use cases are early and low-stakes: rehearsing a discussion guide, pressure-testing a document, surfacing gaps in existing research, and keeping known customer perspectives present in a conversation. Ground them in your own evidence, require a citation for every meaningful claim, and do not let the synthetic user become the decision-maker.

What is a synthetic user?

A synthetic user is a profile generated by a large language model that responds as if it were a real user.

What makes one useful is not how convincingly it sounds human. It is where the profile comes from.

A grounded synthetic user is built from evidence you already have, such as interview transcripts, survey responses, support tickets, sales calls, and product research. Its attributes and responses can be traced back to that source material, and where the evidence runs out, the profile should be able to say so.

An ungrounded synthetic user starts with a description such as: act as a 34-year-old product manager at a mid-market fintech company.

A general-purpose model can generate a plausible response from there. But that answer reflects patterns in the model's training data, not evidence from your own customers.

Both get called synthetic users. They are not the same thing.

The practical difference shows up quickly in language. A general-purpose model tends to produce polished, generic answers. A synthetic user grounded in your own research repository can reflect the language your customers actually use, including the different words roles or segments use to describe the same problem.

Synthetic personas, users, panels, and digital twins

Several terms get used interchangeably in this space, which makes conversations about synthetic research harder than they need to be.

Synthetic persona

A synthetic persona is an archetype grounded in research evidence. It represents a type of user rather than one specific individual: a senior user experience researcher at a business-to-business software company with 500 to 5,000 employees.

Synthetic user

A synthetic user is a more specific individual with attributes, context, and a voice. You might sample one from a broader persona or build one from research on a real person: Sarah, 38, lead researcher at a 2,000-person software company who is frustrated by tool sprawl.

Synthetic panel

A synthetic panel is a group of synthetic users completing the same study.

Instead of asking one simulated participant what they think, you might run ten synthetic users through the same survey or concept test and compare the responses.

Digital twin

A digital twin is a longitudinal model of one specific real person, built from repeated research and other data about that individual. It is narrower than a persona or panel, but potentially much deeper.

Why “how accurate are synthetic users?” is the wrong question

Accuracy sounds like the obvious metric. The problem is that you need a stable ground truth to measure it against, and human responses are not perfectly stable.

Great Question originally designed a head-to-head experiment to compare synthetic and real participants. The plan was straightforward: give a synthetic panel and a matched group of real senior user experience researchers the same concepts, documents, and questions, then compare the themes each group produced.

The problem appeared as soon as the panels disagreed.

If the synthetic users raised an issue the humans did not, was that a useful pattern the humans happened to miss, or was it fabrication? If the real participants raised something the synthetic panel missed, was that a failure of the model or simply an unusual experience from one participant?

The comparison could show differences, but not which side represented the truth.

Sentiment and idea overlap do not completely solve the problem either. You can create scoring rules to measure how similar two sets of answers are, but those rules introduce judgment about what counts as the same idea.

A cleaner question: does this synthetic user reproduce the structure of the evidence it was built from well enough for the job we are asking it to do?

That can be evaluated more usefully through evidence quality, reproducibility, citations, and the stakes of the decision.

What the widely quoted 85% accuracy figure actually means

One of the most frequently cited pieces of evidence for synthetic users comes from research involving 1,052 people and their AI-generated replicas. The headline number is often presented as 85% accuracy.

The underlying numbers tell a more useful story:

  • Humans reproduced their own prior answers about 81% of the time when asked again later.
  • Synthetic replicas reproduced the original answers about 68% of the time.
  • The 85% figure comes from dividing the synthetic score by the human self-consistency score.

That normalization is a reasonable attempt to account for the fact that humans change their own answers.

But “85% of human performance” and “matched the original answer 68% of the time” create very different intuitions about reliability. That distinction matters when deciding what you are willing to trust a synthetic user to do.

There is another problem too: synthetic-user research does not yet use one settled methodology. Different studies build profiles differently, use different source material, test different tasks, and score outputs in different ways. Comparing one headline number with another can make the evidence look more standardized than it is.

The safest interpretation is directional. Synthetic users may help you identify potential preferences, objections, or gaps worth investigating. They are much less suited to telling you the magnitude of an effect or replacing research where the consequences of being wrong are high.

Four ways to build a synthetic user

There is no single synthetic-user workflow. At least four distinct approaches are useful, and each solves a different problem.

1. Build a digital twin

Start with one real person you have unusually rich research on. Remove identifying details where appropriate, then ground the profile in everything relevant you know about that individual:

  • Interview transcripts
  • Product usage
  • Support interactions
  • Account notes
  • Sales conversations
  • Repeated research sessions

The advantage is depth. A digital twin can preserve the perspective of a key customer or design partner when that person cannot be in every internal conversation.

The limitation is equally important: it is one person. A deep model of one customer should not be treated as evidence for an entire segment.

2. Build a segment-based synthetic user

Instead of one participant, combine evidence from several people who share meaningful characteristics.

The goal is not simply to average their demographics. Build the profile around evidence that matters to the research question:

  • Goals and jobs
  • Language and terminology
  • Behaviors
  • Pain points
  • Usage patterns
  • Validated themes

Start with an evidence audit. Do you actually have enough real research on this segment to build something worth using? If the repository contains only a couple of relevant sessions, the answer may be that you need more research, not that the model should fill in the gaps.

This approach can be particularly useful when you want to keep distinct segments present in internal conversations. You might build separate synthetic users for power users, casual users, and churned customers, then run the same document or concept past all three. It pairs naturally with research-backed personas.

The trade-off is aggregation. Combining people into one profile smooths out some of the quirks and contradictions that make real customers messy.

3. Build a synthetic panel

A synthetic panel samples several synthetic users and runs them through the same study. This is closer to traditional panel research than asking one persona a question repeatedly. It can be useful for:

  • Dry-running a survey
  • Finding ambiguous questions
  • Testing whether different segments interpret an item differently
  • Generating hypotheses before recruiting real participants

The quality of the panel is limited by the quality and diversity of the underlying data. A synthetic panel that looks varied on paper but comes from thin or homogeneous source material can create more confidence than the evidence deserves.

When you are ready to test those questions with actual customers, Great Question's research panel lets teams recruit from their own participant base rather than treating synthetic and human research as competing systems.

4. Use live retrieval

The fourth approach does not store a persona at all. Instead, a system queries your repository when you ask a question, retrieves the most relevant evidence, and generates a response from that material.

This can be useful when you want reactions to a specific document or question without maintaining a library of fixed profiles. It can also keep answers current because the system can work from the latest available evidence.

The trade-off is consistency. Two queries against the same repository may retrieve slightly different evidence and produce slightly different answers. That variability matters more when the synthetic user is becoming a shared reference point for the team.

Which synthetic-user approach should you use?

Start with what you are trying to represent.

  • One specific person's perspective: build a digital twin.
  • A recurring customer segment: build a segment-based synthetic user.
  • A spread of responses: use a synthetic panel.
  • A current question or document: use live retrieval.
  • A high-stakes go or no-go decision: recruit real people.

If your first instinct is to create a synthetic user before asking which of those problems you actually have, the tool may be driving the research question rather than the other way around.

Why synthetic users need a research repository

Synthetic users work much better when the source material lives in a structured UX research repository rather than a folder of transcripts. The reason is technical.

Long transcripts contain a lot of information, and large language models do not always use every part of a long input equally well. Research on long-context models has shown that information buried in the middle of large inputs can be harder for models to retrieve effectively, a problem often described as lost in the middle.

Simply pasting dozens of transcripts into a prompt does not guarantee the right evidence will shape the answer.

A stronger system uses retrieval-augmented generation. Instead of asking the model to reason over everything at once, the system retrieves the pieces of evidence most relevant to the question and gives the model a smaller, more focused body of material to work with.

The infrastructure underneath matters.

  • Hybrid search. Combining keyword and semantic retrieval helps connect claims to what customers actually said rather than something merely similar.
  • Filtering. The model should reason only over material relevant to the segment, date range, or research question.
  • Structured metadata. Roles, segments, dates, accounts, research methods, and other fields help define which evidence belongs in an answer.
  • Validated insights. A curated layer of findings gives the system research conclusions that have already been reviewed alongside the raw data.
  • Resolvable citations. Every meaningful claim should lead back to the source.

That last piece is what turns a fluent answer into something a researcher can audit.

How reproducible are synthetic users?

Great Question tested reproducibility by building profiles from the same small survey five separate times using the same underlying data and prompts.

Some things stayed remarkably stable. The strongest segment appeared across all five runs. Distinctive verbatim phrases were repeatedly attributed to the correct respondents. Participants who were clearly outside the target segment were consistently excluded.

Other things moved. The number of profiles generated varied. Confidence ratings on weaker clusters changed between runs. The thinnest segment fractured differently depending on how aggressively each run grouped people together.

That pattern is more useful than a blanket statement that synthetic users are either reproducible or not.

Strong evidence tends to produce stable structure. Thin evidence produces plausible variation. The dangerous part is that the output itself does not tell you which one you are looking at.

Live retrieval has a similar property. The broad themes may stay consistent while wording, examples, and retrieved evidence shift between runs.

For low-stakes exploration, that may be completely acceptable. For something teams are going to use repeatedly, consider maintaining a saved synthetic persona or user as the shared reference point and using live retrieval for one-off questions around it.

How to build a synthetic user that holds up

Start with research you own

Use evidence from real customer interactions. Interview transcripts are especially useful because they preserve language and context. Surveys can add breadth. Support and sales conversations can add recurring operational issues and buying context.

Great Question's user interview guide and survey research guide cover those methods if you need to strengthen the source research first.

Define the population before building

Decide who belongs in the segment before you start clustering people. Ambiguous inclusion criteria are one of the easiest places for two runs to produce different results.

A screener survey can also help create cleaner segmentation in the research you collect now, which makes future synthetic profiles easier to ground.

Set an evidence threshold

Decide what counts as strong, moderate, or thin evidence before looking at the output. The exact numbers matter less than applying the rule consistently and exposing the evidence count alongside the claim.

If a theme appears across many sessions, say so. If it appears in only one or two, label it as a signal rather than quietly giving it equal weight.

Require citations

Every meaningful claim should point back to the research supporting it. No citation should mean no confident claim.

If the model cannot find evidence, it should say the repository does not contain enough information rather than completing the picture itself.

Preserve verbatim language

Direct customer language is one of the strongest pieces of grounding available. It also makes it easier to see whether the synthetic user is reflecting actual research or slipping toward generic model-generated phrasing.

Include gaps

A useful synthetic profile does not only tell you what you know. It should also tell you what the source research does not cover. Those gaps are not a weakness in the output. They are a prioritized list of questions for your next round of customer research.

Build it more than once

Run the same build several times and compare the results. What stays stable deserves more confidence. What changes deserves scrutiny. That reproducibility check can tell you more than how polished the profile sounds.

How to tell whether a synthetic-user answer is trustworthy

You do not need a formal validation study every time you ask a synthetic user a question. A few simple checks go a long way.

Ask for the source

Every claim should identify the research it came from. If there is no source, treat it as speculation.

Ask the question more than once

Run the same question in separate sessions. If the answer changes substantially, the model may be improvising around thin evidence.

If the same structure keeps coming back from different runs and citations, you have more reason to treat it as a genuine pattern in the source data.

Test it against something you already know

Ask about a finding from a recent study. Getting the answer right does not prove the entire synthetic user is valid. Getting a known finding badly wrong tells you immediately that something is off in the grounding or retrieval.

Ask what it doesn't know

A grounded synthetic user should be able to identify gaps. A system that always produces a confident answer, regardless of what the source material contains, is hiding uncertainty rather than managing it.

The synthetic-user uncanny valley

Grounding synthetic users in real customer language makes them much more convincing. That creates another problem: convincing is not the same as true.

One company experimenting with synthetic personas found that early versions built from persona summaries sounded generic. Adding real interview language made different roles sound noticeably more distinct. That was an improvement, but it also made the personas easier to over-trust.

Real people contradict themselves. They change their minds. They misunderstand questions. They have bad days. They care intensely about one issue and barely notice another. Synthetic users tend to smooth away some of that messiness.

The more human a synthetic user sounds, the easier it becomes to forget that you are interacting with a model of the research rather than another participant.

That is why trust but verify is not enough on its own. You need structural guardrails such as citations, evidence counts, confidence labels, and clear boundaries around where synthetic users are allowed to influence decisions.

Where synthetic users fit

A useful governance model is to divide use cases by the cost of being wrong, not simply by research method.

Low-stakes, early-stage work

Synthetic users are well suited to:

  • Brainstorming
  • Pressure-testing early concepts
  • Finding obvious weaknesses in a document
  • Generating questions worth asking real customers
  • Rehearsing a study before launch

The consequence of a poor synthetic answer is usually another discarded idea.

Mid-stage work

Use synthetic users alongside real research. Examples include:

  • Reviewing a discussion guide before interviews
  • Testing whether a concept has obvious problems
  • Keeping a difficult-to-recruit segment present in early conversations
  • Comparing reactions across synthetic segments before deciding what to test with people

If the goal is to expand the number of actual customer conversations you can run, rather than simulate them, Great Question's AI-moderated interviews offer another route. AI handles the moderation while the evidence still comes from real participants.

Strategic and high-stakes work

Use real participants. If the decision is expensive, difficult to reverse, emotionally sensitive, or consequential to customers, synthetic users should not be the deciding evidence.

The same applies when you need magnitude rather than direction. Synthetic users may help you generate hypotheses. They cannot reliably tell you that a specific percentage of your market believes something or that an emotional reaction is representative.

The citation test

If you keep one operational rule from this guide, make it this: no citation, no claim.

Citations are what separate a grounded synthetic user from a plausible character. Every meaningful assertion should connect back to the interview, survey response, support conversation, or other customer evidence behind it. Anything without support should be labeled as low confidence or unknown.

This matters because the most dangerous failure mode is not always an obvious hallucination. It is confident thinness: a polished answer built from weak evidence that sounds exactly like one built from strong evidence.

A good AI analysis and synthesis workflow keeps the claim and source connected so a researcher can inspect the evidence rather than trusting the summary.

What synthetic users change about the research role

There is a larger implication behind all of this.

Traditional research often works one question at a time. A study is designed around a decision, participants are recruited, findings are synthesized, and the team moves on.

Synthetic users make the accumulated evidence itself more reusable. That shifts some of the researcher's value toward building and maintaining the evidence base that future questions will draw from.

The repository stops being administrative infrastructure and becomes part of the research product. That means maintaining clean metadata, preserving source material, curating validated insights, filling meaningful gaps in customer evidence, setting governance around AI use, and making uncertainty visible.

Great Question's AI features, MCP integration, and research repository are built around that broader idea: research as a reusable evidence system rather than a stack of individual reports. The AI-native research guide covers the wider shift.

It also leaves an important question open. What does the research role become when more people interact with customer evidence through AI instead of opening a research report themselves?

The bottom line

Synthetic users are most useful when you treat them as a rehearsal and synthesis tool, not a replacement participant.

They cannot be reduced to one clean universal accuracy score because the human baseline itself moves. What you can evaluate is whether the profile is grounded, whether its claims are traceable, whether its strongest patterns reproduce across runs, and whether you're using it in a situation where being wrong is cheap enough.

Use synthetic users to sharpen a discussion guide before a real interview. Use them to pressure-test a document. Use them to keep existing customer evidence present while your team debates an idea.

Do not use them to avoid talking to people. The better synthetic users get at sounding human, the more important that distinction becomes.

Book a demo

Frequently asked questions

What is a synthetic user?

A synthetic user is an AI-generated profile that stands in for a real research participant. Grounded synthetic users are built from research you already own, with claims tied back to real source material. Ungrounded synthetic users are generated primarily from a prompt and the model's existing knowledge.

How accurate are synthetic users?

There is no single meaningful accuracy figure. Human answers themselves change over time, which makes the baseline unstable. In one widely cited study, synthetic replicas matched original responses about 68% of the time, while people reproduced their own original answers about 81% of the time. The often-cited 85% figure normalizes synthetic performance against that human self-consistency rate.

How do you build a synthetic user?

Four common approaches are a digital twin of one real person, a segment-based synthetic user built from multiple customers, a synthetic panel sampled from broader data, or live retrieval that queries your research repository in response to a specific question.

Can synthetic users replace real user research?

No. Synthetic users are strongest for early exploration, rehearsal, and pressure-testing. Real participants should remain the source of evidence for high-stakes decisions and for questions where new or unexpected customer behavior matters.

What is the difference between a synthetic persona and a synthetic user?

A synthetic persona describes an archetype or type of user. A synthetic user represents a more specific individual with a defined context and voice, either sampled from that persona or grounded in research on a particular person.

How much data do you need to build a synthetic user?

There is no universal minimum. The more evidence you have on a clearly defined segment, the more stable the profile is likely to be. Thin evidence should be labeled as a signal or gap rather than presented with the same confidence as a pattern supported across many sessions.

Are synthetic users reproducible?

They can be more reproducible where the underlying evidence is strong. Great Question's repeated-build experiment found that the strongest patterns and verbatim evidence stayed stable while weaker clusters, profile counts, and confidence ratings moved between runs.

When should you not use synthetic users?

Do not rely on synthetic users for high-stakes or irreversible decisions, emotionally sensitive questions, estimates of magnitude, or situations where talking to real participants is both feasible and important to the decision.

Share