Most “digital twins” in research are synthetic users with a better name

Synthetic users vs. digital twins
✓ how to tell which one you're being sold
Ask a synthetic user a question and you get an answer that's like someone's. Ask a digital twin and you're supposed to get an answer that is someone's. A real person, with a name, who you could go and ask yourself.
That's the whole difference. The trouble is that the market has decided “digital twin” sounds more expensive than “synthetic user,” so the labels have drifted. A persona prompt with a demographic paragraph gets sold as a twin. Simile, a company whose research roots are in agents built from interviews with real people, is described by TechCrunch as a “synthetic-user startup”. If you buy or run research, you need to be able to tell them apart, because they fail in different ways.

Digital twins came from engineering, and meant one specific thing
The digital twin isn't an AI idea. Michael Grieves sketched it in 2002 at the University of Michigan as a model for product lifecycle management: a physical thing, a virtual copy of it, and data flowing between the two. John Vickers at NASA gave it the name around 2010 (Grieves & Vickers, 2017).
NASA and the US Air Force then gave the cleanest definition we have. In Glaessgen & Stargel (2012), a twin is a simulation of an as-built vehicle that uses sensor updates and fleet history to “mirror the life of its corresponding flying twin.” Not a model of an F-15. A model of one particular F-15, tail number and all, updated every time that airframe flies.
Manufacturing researchers later tightened it further. Kritzinger et al. (2018) split the idea three ways by how data moves:
- Digital model. Data moves between the real thing and the model by hand, if at all.
- Digital shadow. Data flows automatically from the real thing into the model, one way.
- Digital twin. Data flows automatically in both directions.
Hold onto the first of those. It comes back later.
When health researchers started applying twins to people, the ethics showed up straight away. Bruynseels, Santoni de Sio & van den Hoven (2018) warned that patterns across a population of personal twins could drive “segmentation and discrimination.” A model of you, specifically, carries obligations a model of “people like you” doesn't.
Synthetic users grew out of survey simulation
Personas and user models are old ideas, but today's synthetic-user debate draws most heavily on a different line of work: social scientists asking whether a language model could stand in for a survey panel.
Argyle et al. (2023) conditioned GPT-3 on the demographic backstories of real survey respondents and called the output “silicon samples.” Their key idea was algorithmic fidelity: whether the model's answers for a subgroup match the real subgroup's answers. Horton (2023) re-ran classic behavioural economics experiments on “Homo Silicus” and got qualitatively similar results. Aher, Arriaga & Kalai (2023) replicated three of four classic human-subject studies, including Milgram.
Notice what all of these are measuring. Not “did the model predict what Maria would say.” They're asking whether the distribution looks right. A synthetic user usually stands for a segment rather than a named person. It can still be built from real participants' data, and it still raises consent questions, but no individual can check whether it got them right.
Synthetic user | Digital twin | |
|---|---|---|
| Represents | Synthetic userA segment or archetype | Digital twinOne identifiable person |
| Built from | Synthetic userA persona prompt, or research about a segment | Digital twinThat person's own data |
| Validated by | Synthetic userDo the aggregate answers match the real group? | Digital twinDoes it predict this person's next answer? |
| Connection to reality | Synthetic userVaries, from an ungrounded persona to a segment built on real research | Digital twinOngoing, refreshed as the person changes |
| Rights and consent | Synthetic userBias, representation, provenance and consent for reuse | Digital twinAll of that, plus likeness, withdrawal and payment |
The number that separates them
The clearest early evidence for why grounding matters comes from a study that compared person-specific agents with demographic and persona baselines.
Park et al. (2024) interviewed 1,052 Americans for two hours each, then built an agent for every one of them from their own interview transcript. On the General Social Survey, those agents predicted each person's answers 85% as accurately as the people themselves reproduced their own answers two weeks later.
Then they ran the control that matters. Agents given only the person's demographics scored 71%. Agents given a written persona paragraph scored 70%.
That's roughly fifteen points from giving the agent one real person's own data. The interview format isn't the magic ingredient, though. A later revision of the paper reports that agents built from the person's survey answers scored 82%, close to the 83% for interview-grounded agents, against 74% for demographics only. What moves the number is data about that specific person.
In Kritzinger's terms, those agents are digital models of one person: built by hand from their data and frozen at interview time. Nobody updates them when the person changes their mind, changes jobs or has a kid. They're a long way past a persona prompt and still two steps short of a twin.
Twins aren't magic either
If you stopped reading at 85% you'd think the problem was solved. It isn't, and these twins were also built once, not kept in sync.
Columbia's team built Twin-2K-500, a public dataset of 2,058 people who each answered more than 500 questions across four survey waves, so twins could be built and tested properly. Then they ran a mega-study: 19 preregistered studies, 164 outcomes. The twins' predictions were only modestly more accurate than a base model that knew nothing about the individual, and across outcomes the average correlation between twin and human answers was about 0.20. The latest version of the paper is titled, perfectly, Digital Twins as Funhouse Mirrors: Five Key Distortions. The five: insufficient individuation, stereotyping, representation bias, ideological bias and hyper-rationality.
That last one has shown up since the start. Aher et al. found a “hyper-accuracy distortion”: on estimation questions, models were far more accurate than any real crowd. Real people are messier than the models want them to be.
And synthetic users fail in predictable ways
The case against using synthetic users as a replacement for participants is now well documented.
- They flatten groups. Wang, Morgenstern & Dickerson (2025) compared four LLMs against 3,200 real participants across 16 identities and found the models systematically misportrayed and flattened identity groups.
- They hold somebody's opinions, just not your users'. Santurkar et al. (2023) found the gap between model opinions and US demographic groups was “on par with the Democrat-Republican divide on climate change.”
- They're too tidy and they drift. Bisbee et al. (2024) found synthetic survey answers had less variance than real ones, shifted with small prompt changes, and changed over three months as the underlying model was updated.
- They can get the sign wrong. The updated version of Brand, Israeli & Ngwe, one of the early optimistic market research papers, now describes LLM willingness-to-pay estimates as “sometimes comparable… often inaccurate and in some cases wrong-signed.”
- Qualitative researchers see the gaps. In Kapania et al. (2025), 19 qualitative researchers saw some uses but flagged missing consent, thin context and the risk of synthetic stories being mistaken for participant data. Agnew et al. (2024) argue that AI surrogates are at odds with the reasons we do participant research in the first place.
Nielsen Norman Group's practitioner verdict is about right: useful for desk research and generating hypotheses, “but not for final decision-making” (Rosala & Moran). Their sharper line: “UX without real-user research isn't UX.”
How to tell what you're actually being sold
Investors are paying attention. Simile, co-founded by Joon Sung Park of the 1,000-people study, raised $100M in February and $200M at a $2B valuation five months later. Aaru is reported as an AI synthetic-research startup, and Electric Twin describes its product as synthetic audiences built from customer data. Most research teams are now being pitched some version of this. Here's the test I'd run on any of it, including anything we build.
- Can you name the person it's a twin of? If not, it's a synthetic user. That's fine. Price it and use it like one.
- What was it built from? A persona paragraph, a segment of real research, or one person's own words?
- When did it last sync with that person? If the answer is “at creation,” it's a snapshot, not a twin.
- How is it validated, and against whom? “Our outputs look realistic” isn't validation. Ask for the comparison against real answers from the real people.
- Did the person agree to be twinned, and can they take it back? If a vendor can't answer this one, treat it as a governance risk, not a twin.
ESOMAR's 20 questions for buyers of AI-based services is worth having open in the same tab.
Where we land
Ungrounded persona prompts are easy to ship. Anyone with an API key and a template can make one. Synthetic audiences grounded in real research take more work, and their value depends on where the data came from, how they're validated and what you let them decide. Either way, they're good for what the evidence says they're good for: stress-testing a discussion guide, drafting hypotheses, finding the obvious holes before you spend a participant's time.
Digital twins are a different kind of product, and they can't be built from a prompt. They need real people who've agreed to take part, a record of what those people actually said, and a way to keep checking the twin against the person it claims to be. That last part is the one that matters most. A twin is the only kind of simulated participant you can audit: ask it a question, ask the real person the same question, and keep score.
It's been awesome to see some of our customers experiment with these technologies before we've rolled out native features to support it. We have customers using their existing research repository to create synthetic users for testing their study plans, screeners and interview guides. We also have some taking it further, running “interviews” with these “users.” I'm not sure how comfortable I am with that approach, but I think these are all essential steps on the path to figuring out how best to apply these new capabilities.
So when someone shows you a “twin,” ask who it's a twin of. If there's no name, there's no twin.
Sources
- Grieves, M. & Vickers, J. (2017). Digital Twin: Mitigating Unpredictable, Undesirable Emergent Behavior in Complex Systems. Transdisciplinary Perspectives on Complex Systems, Springer.
- Glaessgen, E. H. & Stargel, D. S. (2012). The Digital Twin Paradigm for Future NASA and U.S. Air Force Vehicles. 53rd AIAA SDM Conference.
- Kritzinger, W. et al. (2018). Digital Twin in manufacturing: A categorical literature review and classification. IFAC-PapersOnLine 51(11).
- Bruynseels, K., Santoni de Sio, F. & van den Hoven, J. (2018). Digital Twins in Health Care: Ethical Implications of an Emerging Engineering Paradigm. Frontiers in Genetics 9:31.
- Argyle, L. P. et al. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31(3).
- Horton, J. J. (2023; revised 2026 with Filippas, A. & Manning, B. S.). Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? NBER Working Paper 31122.
- Aher, G. V., Arriaga, R. I. & Kalai, A. T. (2023). Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. ICML.
- Park, J. S. et al. (2024). Generative Agent Simulations of 1,000 People. arXiv 2411.10109v1. Revised in v3 (June 2026) as “LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.”
- Toubia, O. et al. (2025). Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions. arXiv 2505.17479.
- Peng, T. et al. (2025). A Mega-Study of Digital Twins Reveals Strengths, Weaknesses and Opportunities for Further Improvement. arXiv 2509.19088 (v1). Latest version titled “Digital Twins as Funhouse Mirrors: Five Key Distortions.”
- Wang, A., Morgenstern, J. & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence 7.
- Santurkar, S. et al. (2023). Whose Opinions Do Language Models Reflect? ICML.
- Bisbee, J. et al. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models Political Analysis 32(4).
- Brand, J., Israeli, A. & Ngwe, D. Using LLMs for Market Research. HBS Working Paper 23-062 (first circulated in 2023 as “Using GPT for Market Research”).
- Kapania, S. et al. (2025). “Simulacrum of Stories”: Examining Large Language Models as Qualitative Research Participants. CHI.
- Agnew, W. et al. (2024). The Illusion of Artificial Inclusion. CHI.
- Rosala, M. & Moran, K. (2024, updated 2025). Synthetic Users: If, When, and How to Use AI-Generated “Research”. Nielsen Norman Group.
- ESOMAR (2024). 20 Questions to Help Buyers of AI-Based Services.




