Launching soon on the Shopify App Store

testfeed

synthetic respondents

Synthetic Respondents: How Accurate Are They vs Human Panels?

Synthetic Respondents: How Accurate Are They vs Human Panels?

Ask how accurate synthetic respondents are and you get two answers. Vendors quote 90 per cent. Sceptics call it a horoscope with a bigger vocabulary. The useful answer sits between them, and where exactly it sits decides whether this technology saves you money or quietly wastes it.

I build a tool that uses this method, so I will be precise about where it works and where it does not. The fastest way to waste money here is to trust the average and forget the rest.

What synthetic respondents actually are

A synthetic respondent is an AI-generated answer standing in for a real survey participant. You describe a shopper, their age, income, habits and what they buy, and a large language model answers your questions as it predicts that person would. Run a few hundred at once and you have a synthetic panel, sometimes called silicon sampling. Nobody is recruited. No survey goes out. You get numbers back in minutes for the price of the compute. They are the most-debated half of AI market research: the part that skips real people rather than just analysing them.

That is the appeal, and it is real: fast, cheap, available at 2am, able to stand in for any demographic you can describe. Convenience was never the question. The question is whether the answers are true enough to act on. So: how accurate are synthetic respondents, actually?

How accurate are synthetic respondents, really?

Start with the strongest evidence in the method’s favour, because it is better than sceptics allow.

In 2024 a Stanford team built generative agents of 1,052 real people, interviewing each for two hours and using the transcript to prime a model to answer as them. Then they checked the agents against the real people. On the General Social Survey, the agents reproduced a person’s answers about 85 per cent as accurately as that same person reproduced their own answers two weeks later.

Read that benchmark twice, because it is the clever part. They did not measure the AI against perfect human consistency. They measured it against how much a real person disagrees with themselves a fortnight on, which is more than most of us like to admit. Against that honest bar, 85 per cent is a serious result. It also quietly reframes the whole debate: humans are not a stable gold standard. Ask the same person the same attitude question two weeks apart and they wobble. “As accurate as a human panel” was always a lower bar than the phrase makes it sound.

The academic case goes back further. John Horton and colleagues at NBER showed in Homo Silicus that if you give a language model an endowment, a preference and a scenario, it reproduces the qualitative findings of classic behavioural-economics experiments, the fairness and framing effects economists spent decades documenting. Lisa Argyle’s team coined the phrase “algorithmic fidelity” for the same effect in surveys: condition the model on real demographic backstories and its answers track the response patterns of actual subgroups. Notice what Argyle’s group actually recommend doing with this. They suggest using it to trial a survey on simulated people before fielding it on real ones. A rehearsal, not a replacement.

So the aggregate story is genuinely good. The trouble starts when you zoom in.

Where synthetic respondents break

The aggregate is where the method looks best. Focus on the things a brand actually pays research to learn and the accuracy falls away in a pattern that is now well documented.

The clearest problem is that simulated answers are too tidy. When Bisbee and colleagues had a model adopt personas and rate social groups, the averages matched a benchmark survey closely and almost nothing else did. The synthetic answers showed far less variation than real people, the correlations between variables came out different, and the model exaggerated how divided and how certain people were. A synthetic panel can nail the headline number and be wrong about the shape of opinion underneath it, which is usually the part you were paying to see.

The second problem is that “any demographic you can describe” is not the same as “any demographic the model represents well”. Stanford’s Whose Opinions Do Language Models Reflect? found models are meaningfully misaligned with several real US groups, with people over 65 among the worst served, and steering the model towards a group did not reliably close the gap. If your buyer sits in a segment the model is weak on, a synthetic panel will be confidently unrepresentative and give you no warning it is off.

The third problem is the one that costs the most. Synthetic respondents are best at repeating what the world already knows and shakiest on anything new, which is exactly when you reach for research in the first place. Harvard Business School researchers used a model to estimate willingness to pay for household products. For features the model had effectively seen priced before, it matched real human willingness to pay respectably. For new features it needed tuning against real human data first, and its estimates were sometimes not just off but pointing the wrong way. A tool that is reliable on the familiar and unreliable on the novel is backwards from what you want, because you rarely commission research to confirm the obvious.

Compared to what? Human panels are not truth either

Before this reads as a takedown, the fair counterweight. Human panels are not a clean benchmark. Pew benchmarked online samples against known population values and found opt-in panels were off by an average of 5.8 percentage points, against 2.6 for probability-based samples. Separately it found a real share of online respondents are simply bogus, speeding through for the incentive. So the choice is not flawless humans versus flawed AI. It is two imperfect signals, and the skill is knowing which flaw you can live with for the decision in front of you.

So when can you trust a synthetic panel?

Put the evidence together and a rule falls out. Trust synthetic respondents in the aggregate, for ranking and screening, on familiar ground. Distrust them for the spread, the subgroups and the new. Here is how that maps to real decisions.

The decisionTrust a synthetic panel?Why
Screen 20 concepts down to the 3 worth testingYes, as a directional filterRanking in the aggregate is its strength
Pre-test and debug a survey before you field itYesThe use its own inventors recommend
Get a rough read on a claim or a headlineYes, directionallyFine for triage, not for the final wording
Size how two customer segments differNoIt flattens the variation between people
Set a price to the poundNoWeakest exactly on new, specific numbers
Give the final go or no-go on a launchNoToo expensive to rest on a directional signal
Judge taste, texture or smellNoNo model tastes anything

The top rows are cheap wins you should probably already be taking. The bottom rows are where brands get burned treating a simulation as a verdict.

How to vet a synthetic-respondent vendor

Because the accuracy depends so heavily on the specifics, the marketing number is close to meaningless on its own. “90 per cent accurate” against what, on which categories, measured how? Two industry bodies have turned that into questions you can ask.

ESOMAR, the global standards body for research, publishes 20 Questions to Help Buyers of AI-Based Services. It is the quickest way to separate a serious vendor from a confident one. And the Market Research Society’s Delphi group, in its report on synthetic respondents, makes the point that matters most: to produce useful synthetic data you still need genuinely good human data underneath it, and the profession has yet to agree when synthetic evidence is solid enough to be decision-grade. Until it does, the burden of proof sits with the vendor, not with you.

Four questions do most of the work. What real human data is the model validated against, and how recently? Does it show you the spread of answers or only the average, because only ever showing the average is the tell of a weak tool? Which product categories has it been tested in, and is yours one of them? And what do they recommend you do with a result, because an honest vendor tells you to validate the winners with real people, and a weak one tells you to ship.

Where TestFeed fits

Since I run one of these tools, hold it to everything above. TestFeed lets you put a product, a pack, an ad, a claim, a name, an in-context price or a concept in front of your target shopper and get back a purchase-intent read, the reasons in shoppers’ own words, and a clear next move, in days rather than weeks. The job it is built for is triage: killing the weak ideas cheaply so your stock, budget and live research go to the ones that earn it. We have worked with challenger brands like Bae Juice and Sol Bevi doing exactly that.

Its limits are the method’s limits, which is why we are upfront about how far the signal goes. It is a pre-spend, directional signal, not a sales forecast or a market size. It does not judge taste, texture or smell. And it does not replace real customers when the decision is big enough to need them. Any synthetic tool that tells you otherwise, mine included, would be overselling.

A workflow that keeps the speed without buying the risk

The brands that get value from synthetic respondents are not the ones who trust them most. They are the ones who put them in the right slot in the process.

Use a synthetic panel first, to test an idea against your audience and rank a pile of concepts, claims or packs cheaply, then cut the obvious losers. Take the survivors to real people: a proper concept test for the ideas worth it, a real pricing method like Gabor-Granger for the actual number, and your own reviews and support tickets, read through a voice-of-customer lens, for the truth you already own. After launch, watch the till, because the only fully honest respondent is a paying customer. Synthetic first for speed, humans second for confidence, the market last for truth.

Frequently asked questions

What are synthetic respondents?

Synthetic respondents are AI-generated answers built to stand in for real survey participants. You describe a type of shopper and a large language model answers your questions as it predicts that person would. A few hundred of them together make a synthetic panel, also called silicon sampling. Nobody is recruited and no survey is fielded, so you get results in minutes rather than weeks.

How accurate are synthetic respondents compared to human panels?

Accurate in the aggregate, unreliable in the detail. The strongest published test, Stanford’s 2024 study of 1,052 people, found AI agents matched individuals’ survey answers about 85 per cent as well as those people matched their own answers two weeks later. But the same body of research shows synthetic answers bunch around the average, understate the differences between customer groups, and struggle most with new products. Trust them for direction, not for the spread or the specifics.

Can synthetic respondents replace human panels?

Not for decisions that are expensive to get wrong. They are a strong first filter for screening and ranking ideas cheaply, and useful for pre-testing a survey before you field it. But for a launch, a reformulation or a national price you still need real people, because that is exactly where simulated respondents are least reliable. The sensible pattern is synthetic first for speed, human panels second for confidence.

What is the difference between synthetic respondents and synthetic panels?

They describe the same method at different scale. A synthetic respondent is one AI-generated answer set standing in for a single participant. Synthetic panels are groups of them assembled to mimic a whole sample of your market. The accuracy questions are identical: both are strong on averages and weak on the variation between people.

When should you use synthetic respondents?

Early and often for low-stakes, high-volume work: screening lots of concepts, pre-testing surveys, getting a rough read on claims, or approximating a hard-to-reach audience you cannot recruit quickly. Avoid them for the final call on anything costly, for sizing how segments differ, for pricing to the pound, and for anything involving taste, texture or smell.

The one rule to take away

A synthetic panel is a filter, not a jury. Used to rank and screen before you spend, it is one of the cheapest research upgrades on offer. Used to make the call on a decision that would hurt to get wrong, it hands you a confident average and hides the risk in the part it cannot see. Put it first in the process, keep real people last, and you get the speed without paying for it in accuracy you did not know you were losing.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first study with you.

Direct install from the Shopify App Store arrives in a few weeks.