We built a synthetic user skill. It's now yours.
Four editions of building and testing, plus one experiment we killed halfway through. This week we open up the skill, give you a decision tree and wrap up our findings.
Catch up · Part 1, 2 & 3
Synthetic users: we're building in public
The vocabulary, the ways to build one, and the priors I took into this experiment.
Four editions · One live experiment
We built in public, start to finish.
- 01
The vocabulary, the ways to build a synthetic version of your customer, and the priors going in.
- 02
The data audit, the rigor framework, and the step-by-step build from 726 transcripts.
- 03
We tried to grade synthetic against real, couldn't, and it changed what we think they're for.
- 04
The skill & wrap-up
Our top findings on synthetic users, when to use them, plus the Claude skill you can run yourself.
Four editions ago, I said most of the internet was arguing about whether synthetic users should exist while almost nobody was building one in public and telling you what they found. That was the whole point of this series: build one out in the open and show you the parts that didn't go to plan.
So before I open up the skill, I want to be honest about where the experiment actually landed.
Where it landed
Part 3 was meant to be the head-to-head: a synthetic panel against a real one, with a clean accuracy number at the end. We killed it. You can't grade a synthetic user's accuracy when you can't grade a human's either, and I wrote up the whole detour in Part 3.
The takeaway was: stop asking how accurate a synthetic user is, and start asking what it's good for and when to trust it. That question shaped the skill I'm shipping today.
My top findings
Between building this on our own data and a steady stream of calls with product and research teams building the same thing, here's what's valuable to know before you start your own experimentation.
01
It lives or dies on your own data
A synthetic user built off a generic "market" model knows nothing about your product or your customers. It's a statistical impression of a demographic an LLM read about online, and it answers like one. The whole value is in grounding it on your own transcripts. If the evidence on a segment isn't there, you've found your next research project, and no amount of clever prompting fills that in.
02
Build from verbatim, never summaries
One team I spoke to built theirs off the study summaries in their repo, and it made things up. You want the skill pulling exact verbatim from the transcripts. The real quotes are what make it echo your customer's language back to you instead of something generic and "clean."
03
A repo beats a pile of transcripts, every time
Dumping 20 or 30 transcripts will somewhat work, but not very well: it loses the middle of a long context and fills the gaps with things that sound right. Point it at an indexed repo where the retrieval is already done, then supplement with a handful of sales calls or support transcripts if you want more dimensions to your synthetic user.
04
Treat it as a sparring partner, not a final insight
Every team I talk to that has experimented with synthetic users, use it to react to concepts and rough prototypes as a fast first pass, before they hand it to research. Nobody's presenting "insights from a synthetic user" as the final finding. It's the dry run, and the real interviews still come after.
05
Watch out for 'clean' answers
Real people say surprising things. A synthetic user, even one built on real verbatims, hands you a smooth, low-variation answer. It's strongest on directional, yes-or-no, good-or-bad questions and weakest on nuance: the emotional, identity-laden ones, and brand-new concept areas where your repo has no current research on. That last case is exactly where it'll confidently invent, which is what the confidence meter in our synthetic user skill is there to catch.
06
Make it argue with you
Because it has so much context on you, it's agreeable to a fault. The sycophancy problem again. So build an adversarial version: instruct it to critique, or ground it in churned and closed-lost interviews so it plays skeptic instead of fan. A synthetic user that pushes back is worth ten that nod along.
07
It goes stale, so let the gaps drive you
The data ages. Refresh it through the MCP when you have new transcripts, and when it flags something it can't answer with confidence, treat that as your next research brief. Go fill it with humans, then fold the new evidence back in to your repo and update your synthetic user.
08
And the one that surprised me most: it's fast
With the MCP wired up, I've built a grounded synthetic user in about 5 minutes. Ask it which customers you have the most data on, get a shortlist, sanity-check it isn't hallucinating, and go.
The decision tree
So which workflow, when?
Across Parts 1 and 2 I built four ways to make a synthetic version of your customer. People keep asking which one is "best." They do different jobs and can be used based on what you need to know.
One more the tree bakes in: if consistency across your team matters, and in enterprise it almost always does, keep one saved synthetic user as the official reference and use live retrieval for one-off questions on top of it. Live retrieval causes drift across synthetic users.
Inside the skill
I've been referencing "the skill" for four editions. It runs entirely on the Great Question MCP, the only source it's allowed to touch, and it does two jobs.
Mode 01 · Default
Point it at the repo, it builds profiles
No artifact needed. It surveys the sessions, groups them into clusters by role, workflow and shared pain, counts the evidence behind each, and hands back reusable profile cards you can summon later.
Mode 02 · On demand
Hand it a PRD, a profile reacts
It pulls the relevant evidence with hybrid search, then answers in the first person as that customer, pushing back where the evidence contradicts the artifact and flagging anything the repo can't speak to.
The design principle behind both, in the skill's own words:
You are not a generic AI persona. You do not extrapolate beyond what the evidence supports. When the evidence isn't there, you say so and flag it as a research opportunity rather than fill the gap.
Everything else is enforcement of that one idea. Five non-negotiable rules run on every response, whatever you ask.
Rule 01
Evidence-backed claims only
Every claim, reaction, concern, or pushback you return must be backed by retrieved evidence from the Great Question MCP. No source, no claim.
Before it includes any claim, it runs a self-check: does the retrieved quote explicitly support this specific point, or am I making an inferential leap? Am I confusing a study title with what a participant actually said? If the answer is "unsure," the claim gets dropped and logged as a gap.
Rule 02
A confidence threshold baked in
The 8-interview bar I've mentioned since Part 2, is hard-coded into the skill. It mirrors the pattern-matching threshold our own research team uses, and it always shows the count, e.g. "Mentioned in 12 of 14 interviews across 3 studies."
That count is the confidence flag Part 3 argued was the most useful thing the tool can give you, the error bar you otherwise can't feel.
Rule 03
Cite every claim, and make the citations resolve
Every point traces back to the session it came from, quoted verbatim, by anonymous speaker handle. The rules are strict on purpose: quotes must be at least three words and copied exactly from the transcript search, badge text has to match the quote, and it's forbidden from ever citing a study summary as if it were something a participant said. No quotation marks, no invented paraphrase dressed up as evidence. Of course, I do encourage you to do your own checks here...
Rule 04
Surface gaps
Every response ends with a Research Gaps section: the topics where the evidence ran out, the claims it would have made with more data, and suggested questions to close the hole. The line in the skill file is blunt:
The gap is valuable, flag them.
Filling these gaps in future will also help you build stronger synthetic users in the future.
Rule 05
Stay in character, stay honest
It speaks in the first person as the customer ("this is the part I'd push back on…") because that's what makes it useful to react to. But voice never overrides evidence. If it can't back something up, it doesn't say it, not even in character. A reaction is still a perspective grounded in a real quote.
Two more things. Anonymisation runs by default: speakers only ever appear as handles like Speaker_abc123_1, never real names, which is what makes the aggregated workflow safe.
And it's honest about non-matches: if it finds nothing relevant, it says so and recommends the research to run, rather than dressing up tangential quotes to look thorough.
That's the whole skill. Its only job is to refuse to say more than the evidence allows, which after Part 3 is exactly what I want it doing.
Live experiment · Part 4 of 4
Free synthetic user skill - now yours!
The synthetic-user skill is live. Point it at your Great Question repo and it'll build profiles on day one. Hand it a PRD and one of them will react. And I want testers. Building this in public only works if the last step is public too. Run it against your own repo and tell me where it breaks, especially where a citation doesn't hold or where it should have flagged a gap and glossed over it instead. A full synthetic user guide is coming soon: the terminology, when to use it, and more. Sign up to our newsletter and we'll send it the day it lands.
If you've followed along since Edition 1, thank you.
Get the blog in your inbox
Latest insights for customer-obsessed product builders. Once a month, no noise.

