We built synthetic participants for a recent project; profiled, sourced, ready to go. But we had to kick one of them out.
Not because the model misbehaved. Because we didn’t have enough real evidence about the person behind it, and a persona built on thin data isn’t a participant. It’s a guess with a name attached.
That decision cost us a voice in the sample. It is also the clearest illustration we have of how this kind of research should be done. Here is the standard we hold ourselves to. Synthetic participants, AI personas built from real supplementary data about real people, are increasingly part of the qualitative toolkit, and increasingly used without guardrails.
It starts with the data, not the model
A synthetic participant is only as good as the material behind it. Before we build any persona, we assess both the quantity and the quality of the supplementary data available for that individual.
Where the evidence is too thin to extrapolate from responsibly, we remove the participant rather than force it. Good synthetic research starts with good data, and it takes discipline to walk away when the data isn’t there.
Some questions should never go to a model
Not every question suits an AI-generated answer, and our discussion guides are built accordingly. Where a topic risks pushing a model beyond its training data, current use cases or a fast-moving competitor landscape, for example, we keep that ground human.
The risk is hallucination dressed up as insight. No client should have to second-guess which is which.
Plausible isn’t the same as true
Every synthetic output is checked back against source material and supplementary data to confirm it is grounded in real evidence. Anything that looks hallucinated is flagged and removed before it reaches analysis.
We also score every synthetic participant for trusted groundedness system out of five, where three is the best case: extrapolation that stays in line with a participant’s established thinking patterns and the evidence we hold on them, rather than pure invention. We monitor those scores across every project as a quality gate, not a formality.
Why we’re telling you about the one we deleted
Synthetic participants aren’t a replacement for human insight, and they aren’t a shortcut. Used well, they are an extension of real evidence: bounded by good data, deployed selectively, and checked relentlessly for groundedness.
The participant we removed is the reason you can trust the ones we kept.