Launching soon on the Shopify App Store

testfeed

ai focus groups

AI Focus Groups: How They Work, and When to Use Them

AI Focus Groups: How They Work, and When to Use Them

Two products both call themselves an AI focus group. One runs interviews with real people and writes up the themes overnight. The other never speaks to a single person. Buy the wrong one for the job in front of you and you will either overpay for depth you did not need, or make a launch decision on opinions that were invented by a model.

That is the problem with the phrase. AI focus groups have become a catch-all for two methods that share almost nothing except the marketing. This guide separates them cleanly, shows how each actually works, lays out what the research says about whether you can trust the output, and gives you a decision table and a worked example so you can pick the right one this afternoon.

What an AI focus group actually is

Start with the thing it is replacing. A traditional focus group puts eight to ten recruited people in a room with a moderator, a discussion guide, and usually a client watching through a mirror. It is good at one thing: hearing real people react to something in their own words, and watching them build on each other. It is also slow and expensive. A well-run online focus group costs roughly $4,000 to $7,000 per group, and a full project across several groups and markets runs anywhere from $24,000 to $90,000, according to cost breakdowns from market research firm Drive Research and Greenbook. Recruitment and scheduling push timelines out to weeks. That cost and lag is the gap every AI tool in this space is trying to close.

They close it in two fundamentally different ways.

The first keeps the humans and swaps out the human labour around them. AI moderates the conversation, asks the follow-ups, transcribes everything, codes the themes, and builds the summary. You are still hearing from real people.

The second removes the humans altogether. AI personas, built to represent a described audience, answer your questions directly. Nobody is interviewed. You are reading a model’s best estimate of how that audience would respond.

Everything downstream, cost, speed, and how much you should trust the result, follows from that single fork. So keep the two apart in your head from the start.

The two types, side by side

AI-moderated focus groups: real people, AI runs the room

Here the AI is the moderator and the analyst, not the respondent. A platform sends your discussion guide to real recruited participants and runs the sessions, often as many one-to-one video or chat interviews happening in parallel rather than one group at a time. The model asks the scripted questions, probes when an answer is thin, and then transcribes, tags, and clusters the responses into themes with supporting quotes.

The advantage is scale and speed on genuine human input. Instead of one group of ten, you can run a hundred conversations in the same window and have the analysis drafted by morning. Tools in this camp include Conveo, Voxpopme, Discuss.io and Remesh. Vendors report large cost and time savings, though those figures are marketing rather than independent measurement, so treat them as directional.

The honest trade-off: parallel one-to-one interviews are not the same as the group dynamic of a classic focus group. You lose the cross-talk, the moment where one person’s throwaway comment sets off three others. For a lot of exploratory work that dynamic is overrated and the loss is worth it. But if the group interaction itself is the point, an AI-moderated format is a different thing wearing the same label.

Synthetic focus groups: AI personas, no people

Here there are no participants. You describe an audience, a 34-year-old time-poor parent who buys mid-market snacks, a lapsed gym-goer, a Shopify store’s repeat buyer, and AI personas generate responses as if they were those people. You can ask them to react to a product, rate three ad concepts, or explain why they would not buy.

This idea has a name and a research lineage. It is a commercial application of what academics call silicon sampling, from the 2023 Political Analysis paper by Argyle and colleagues, which showed that a large language model conditioned on real demographic backstories could reproduce the response patterns of specific human subgroups with what they termed algorithmic fidelity. Tools in this camp include FocusGroups AI, Minds, and others positioning around synthetic audiences and synthetic respondents.

The advantage is that it is close to free, instant, private, and infinitely repeatable. You can test twenty subject lines before lunch and nobody sees your unreleased packaging. The trade-off is the obvious one: there are no real people in the loop, so the quality of the answer depends entirely on how well the model represents your audience. That is exactly where the research gets interesting, and where most articles on this topic go quiet.

What the research actually says about accuracy

This is the question that matters, and it deserves a straight answer rather than a sales pitch in either direction. The evidence cuts both ways.

On the encouraging side, models are better at this than sceptics expect. Beyond the Argyle work on algorithmic fidelity, a Harvard Business School working paper by Brand, Israeli and Ngwe found that willingness-to-pay estimates derived from GPT responses were realistic and broadly comparable to estimates from human studies, including a downward-sloping demand curve, people wanting less at higher prices, which is the sort of basic economic sanity check synthetic data often fails. For average preferences and directional reads, the models hold up.

On the cautionary side, the failures are specific and they matter. The most cited critique is Whose Opinions Do Language Models Reflect? by Santurkar and colleagues, presented at ICML in 2023. Measuring model outputs against real US public opinion across 60 demographic groups, they found substantial misalignment, and crucially that the gap persisted even when the model was explicitly steered towards a group. Older and widowed respondents were among the worst represented. The model has a default voice, and pushing it towards a persona only moves it so far.

The second recurring problem is homogenisation. Synthetic respondents tend to reproduce the average answer reasonably but flatten the spread around it. Minority views get compressed, disagreement gets smoothed over, and the output looks more consensual than a real sample would. Market research agency Verian documents this as one of the most consistent limitations of synthetic samples: the means can look right while the variance, the actual texture of human disagreement, quietly disappears. Our deep-dive on synthetic respondents digs further into how accurate this output really is.

Put the two sides together and you get a usable rule. Synthetic focus groups are reasonable for the centre of the distribution, average preference, rough direction, obvious duds, and unreliable at the edges, minority segments, strong dissent, genuinely novel products the training data never saw. They tell you where the crowd leans. They are poor at telling you who in the crowd will hate it and why.

Cost and speed: the honest comparison

The numbers below are directional. Vendor pricing varies and the AI categories are moving fast, but the relative order is stable.

MethodTypical costTime to insightWho you are hearing from
Traditional focus group$4,000 to $7,000 online per group; $24,000 to $90,000 per full project2 to 6 weeksReal recruited people, live
AI-moderated interviewsA fraction of a traditional project; scales with participant numbers24 to 48 hoursReal recruited people, AI-run
Synthetic focus groupLowest; no participant costMinutes to hoursAI personas, no people

The pattern is clear. As you move down the table you trade human reality for speed and cost. The mistake is treating that as a single spectrum where cheaper is always worse. It is not a spectrum. AI-moderated and synthetic are different tools for different jobs, and the right answer is usually to use more than one.

When to use each

The decision is not “AI or humans”. It is which method fits the stakes of the decision and the kind of answer you need.

Your situationBest method
Final sign-off on a major launch, or a regulated claim you will defendRecruited human panel, traditional or AI-moderated
You need the “why” behind a reaction, with real human depth, faster and cheaper than a panelAI-moderated interviews
Early screening of many variants: ad concepts, names, pack designs, price pointsSynthetic focus group
A quick pre-spend read on demand before you invest in production or mediaSynthetic focus group
Anything involving taste, texture, smell, or physical handlingReal people, in person
A niche or under-represented audience where minority views are the whole pointReal people; synthetic will flatten them

If a wrong answer costs you real money or a compliance problem, keep humans in the loop. If a wrong answer just means you refine and try again, synthetic earns its place by letting you fail cheaply and often.

A worked example

Say you run a challenger snack brand. You have three pack designs for a new flavour, a modest budget, and a launch date. Here is how the methods stack rather than compete.

Start synthetic. Put all three designs in front of AI personas built around your actual buyer and ask for a purchase-intent read and the reasons behind it. In an hour you learn that design two splits the audience, design one is nobody’s favourite but nobody’s enemy, and design three scores highest on intent. That does not settle it, but it cuts three options to two and tells you what to probe. Cost: negligible.

Then go to real people on the finalists. Run AI-moderated interviews with thirty of your target buyers on designs two and three. Now you get the “why” in human words: design three reads as premium to some and expensive to others, and the flavour name confuses a chunk of people. That is the kind of specific, quotable, sometimes contradictory feedback synthetic tends to smooth away. Cost and time: a fraction of a classic project, back in a day or two.

Finally, if this is a national retail launch with real money behind it, validate the winner with a recruited panel or a live group before you commit to the print run. If it is a limited online drop, you may not need to. The stakes set the ceiling.

The point of the example is that the question is never “which one is best”. It is “which one, at which stage”. Synthetic narrows and directs. AI-moderated explains. Human panels confirm the expensive decisions.

Where synthetic testing fits for ecommerce and CPG

For merchants and challenger brands, the honest use of synthetic testing is the top of that funnel: getting a fast, cheap read before you spend, so the expensive research is aimed at the right question. That is the job TestFeed is built for. You can test a product, a pack, an ad, a claim, an in-context price, a concept, a name, a promo, or even demand before you spend real money on it, and get back a purchase-intent score, a clear verdict, the shoppers’ reasons in their own words, and a recommended next move, in days rather than weeks. We have run this with brands like Bae Juice and Clutch Glue.

The framing that keeps it useful is the same one this whole guide argues for. It is a directional signal, not a guarantee. It is a pre-launch, pre-spend read, not a sales forecast or a promise of a number. And it will not tell you whether the product tastes good, because no synthetic method can judge taste, texture or smell. Used that way, as the first filter before you commit budget rather than the last word before you launch, it does exactly what an AI focus group should: it makes failing cheap. For a fuller picture of the category and how the pieces fit, our guide to AI market research covers the wider toolkit.

How to run an AI focus group without fooling yourself

Whichever type you use, a few habits separate a useful read from a confident wrong answer.

Describe the audience specifically. “Our customer” is useless. Age, context, budget, what they currently buy and what they are trading off. The more concrete the brief, the less the model falls back on its default voice, which is the failure mode Santurkar and colleagues documented.

Ask about behaviour and trade-offs, not just opinions. “Would you buy this at $24 instead of your current $18 option, and what would you give up?” beats “Do you like this?” every time, with synthetic personas and real people alike.

Read the spread, not just the average. If a synthetic panel comes back unanimously delighted, be suspicious rather than pleased. Real audiences disagree. Flat consensus is usually the homogenisation effect, not a great idea.

Match the method to the stakes. Validate anything expensive, regulated, or irreversible with real humans. Use synthetic to kill weak ideas and sharpen strong ones before they get near a budget.

And never test the physical with the synthetic. Taste, texture, smell, fit, and feel need a body in the room.

The short version

An AI focus group is two tools wearing one name. AI-moderated interviews give you real human depth at a fraction of the old cost and time, and can genuinely replace a lot of classic qualitative work. Synthetic focus groups give you an instant, near-free directional read that is strong on average preferences and weak on the tails, which makes it a first filter, not a final verdict. The teams who get value from either are the ones who match the method to what the decision can afford to get wrong, and who keep real people in the loop for the calls that count.

Frequently asked questions

What is an AI focus group?

An AI focus group is one of two methods. In an AI-moderated focus group, artificial intelligence runs interviews or group discussions with real people, then transcribes and analyses the responses. In a synthetic focus group, AI personas modelled on a described audience answer the questions instead of real people. The two share a name but produce very different data.

Are AI focus groups accurate?

It depends on the type and the question. AI-moderated interviews collect real human answers, so accuracy is about the quality of the moderation and analysis. Synthetic focus groups rely on a model’s estimate of how an audience would respond. Research shows models reproduce average preferences reasonably well but compress the diversity of real responses and under-represent minority views, so synthetic output is best used as a directional signal rather than a final answer.

How much do AI focus groups cost?

A traditional online focus group typically runs $4,000 to $7,000 per group, and a full project across several groups can reach $24,000 to $90,000. AI-moderated interview platforms cut that sharply by running sessions in parallel and automating analysis. Synthetic focus groups are cheaper again because there are no participants to recruit or pay.

Can AI replace traditional focus groups?

Not entirely. AI-moderated interviews can replace much of what a classic focus group does for exploratory and directional work, faster and at lower cost. Synthetic focus groups work as a first pass before you spend. For high-stakes launch decisions, regulated claims, or anything involving taste, texture or smell, real people are still needed.

What is the difference between AI-moderated and synthetic focus groups?

AI-moderated focus groups use AI to run and analyse conversations with real human participants. Synthetic focus groups use AI personas in place of people, so no human is interviewed. AI-moderated gives you genuine human voices at scale. Synthetic gives you an instant, low-cost estimate you can iterate on, with the caveat that it is a model’s guess.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first study with you.

Direct install from the Shopify App Store arrives in a few weeks.