Launching soon on the Shopify App Store

testfeed

The method

No black boxes.

Synthetic research gets a lot of eye-rolls, and often earns them. So here is how ours actually works, where the genuinely hard problems are, and the line we draw between what we will answer and what we won't. Read it and decide for yourself.

A 1950s research scientist in a white lab coat holding a red clipboard, mid-explanation

Why this is hard

You can't just ask ChatGPT.

Prompted well, a language model can reproduce the rough shape of what a population thinks. That is the promise the whole category is chasing. The naive version, though, fails in ways that are now well documented: ask a model to score something and it agrees too readily, avoids the extremes, and collapses a varied population into one agreeable voice.

The output looks like data and behaves like a guess, which is worse than having nothing, because you will trust it. Everything below is how we handle each of those failures, and where we decided the right answer was to not answer at all.

A row of 1950s metal robots at a research table all holding up the same middling score card, one red, beaming in bland unison

The hard problems

The criticisms are real.

These are the documented weaknesses of using language models as respondents. We do not wave them away. We built around the ones we can and refuse the ones we can't.

They agree too easily.

Models are trained to be agreeable, so a leading question gets you the answer you were hoping for. The field calls it sycophancy.

What we do Simulated shoppers never see a scale or a yes/no. Every answer is open text in their own words, so there is nothing to agree with.

They pile on the middle.

Ask a model to rate something 1 to 5 and it clusters on 3, almost never choosing 1 or 5. A naive survey returns tidy, unrealistic numbers.

What we do We do not ask for numbers. Scores come from the language afterwards, so the real shape of opinion survives.

They flatten to the average.

Left alone, models drift toward one bland majority voice and lose the tails, which makes a synthetic sample look calmer than the real thing.

What we do Audiences come from census data, not invention, and every result shows its consensus and range, never just a mean.

They don't speak for everyone.

A model's default view is misaligned with plenty of real groups, something measured directly against live survey data.

What we do We make no claims at the segment level. Our own testing shows those cuts are unreliable, and we say so.

Personas are brittle.

Change a few words in a persona prompt and the answer moves. Personas offer the feel of rigor without the substance.

What we do We do not sell personas. Simulated shoppers sit on real population structure and run through one fixed, versioned protocol.

They have no senses or memory.

A model cannot taste a drink, has not lived with your brand, and has no real memory of last year. Push it there and it invents.

What we do Those studies are blocked before they run. If we cannot answer it honestly, you cannot run it.

How words become a number

The core is published, not homemade.

Simulated shoppers answer in words. To turn those words into a score we use Semantic Similarity Rating, a method published in 2025 by Maier and colleagues. Each answer is matched against reference statements for every rating level, and the score is read from where it lands. The model is never asked for a number.

We built on published, independent research on purpose, rather than a scoring trick of our own. On its authors' benchmark of 57 product surveys, the method recovered about 0% of human test-retest reliability. Grounding it in census data, gating it, and turning it into a verdict is our own work.

01 The answer

“Looks less sugary than the big brands. I'd try a box if the price was close to what I pay now.”

Simulated shopper №072 · open text

02 Matched to anchors

  • 1 Would not buy
  • 2 Unlikely
  • 3 Might
  • 4 Likely
  • 5 Definitely

Similarity to each level's reference statement. No scale is shown.

03 A distribution

1
2
3
4
5

A probability across the scale, not a single guess.

04 Score & verdict

67 Viable

Combined across 100 shoppers, with a confidence range.

Where the shoppers come from

Census-grounded, pattern-shaped, never invented.

Real population structure

Audiences are built from government census microdata, roughly 100,000 US profiles and a tighter Australian set, carrying real age, location, income and occupation.

Shaped by your store

Your Shopify patterns, from repeat purchase to spend bands to product affinities, shape that base as aggregate context. They shape the audience. They never replace it.

No personal data, structurally

Customers are reduced to anonymized summaries in code, with names, emails and addresses stripped and small groups suppressed. There is no personal data to leak because none is kept.

A 1950s records office: a line of people files out of an open cabinet drawer down a ramp of folders, watched by a clerk in cat-eye glasses holding a red clipboard

Where it works, and where it doesn't

Built for consumer, by design.

Language models are strongest where they have absorbed the most: everyday retail, food, drink, packaging, household goods. The method we build on says the same, that it only holds where the model already knows the domain well. Our strength in consumer is not luck. It is the shape of the tool, confirmed across a lot of categories we tried and chose not to keep.

Where we're confident

  • Consumer, retail, ecommerce and FMCG decisions
  • Purchase intent, our most reliable read
  • Packs, claims, ad concepts and value propositions
  • Bounded price viability, in context
  • Head-to-head reads, and honest ties
  • Directional, pre-launch signal

Where we're not

  • B2B and niche, expertise-dependent categories
  • Taste, smell and texture
  • Open "how much would you pay?"
  • Brand awareness, recall and loyalty
  • Segment-level precision
  • Sales, conversion or market forecasts
A 1950s scientist in a white lab coat gives a thumbs up beside a chart of two near-identical bars marked with a red approval tick

The label on everything

Directional, every time.

Every number we return is a signal for a decision, not a forecast of the market. We would rather be usefully right about the direction than falsely precise about the destination, and we label it that way every time, so you always know exactly what you are holding.

Try it on your own store.

Book a demo and we'll run your first study with you.

Direct install from the Shopify App Store arrives in a few weeks.