Launching soon on the Shopify App Store

testfeed

For Shopify merchants

Test it on 100 shoppers before you spend a dollar.

Run your products, packaging, pricing and ads past simulated shoppers built from census data and your store's own patterns. Get a purchase intent score, a clear verdict and the reasoning behind it before you spend real money.

  • Built for Shopify
  • Grounded in census data
A 1950s woman proudly presenting a Sunny Crisp corn flakes box and a bowl of cereal

Purchase intent

0 Viable

Put to work by

AndieClutchBae JuiceSol BeviElita
Four cheerful people in 1950s dress seated in a row, one writing on a red clipboard

How it works

Four steps. One verdict.

Pick a product.

Install the app and pick any product from your catalog. That's what gets tested.

Write your questions.

Open-ended questions, like “What would stop you buying this?” testfeed blocks question types it can't answer honestly before they run.

Pick your audience.

Simulated shoppers start from US or Australian census data, shaped by your store's aggregate purchase patterns. No personal customer data is used.

Get the results.

A purchase intent score, a verdict, and every shopper's full answer, inside Shopify admin.

Results

What you get back.

Every study returns a purchase intent score with a confidence range, the level of agreement among shoppers, and every individual answer in full. The verdict is one of four bands, based on the score.

Strong Viable Iterate Weak

Ties are reported as ties

testfeed never manufactures a winner. If two variants perform the same, it tells you, and you can stop debating.

Pack test · Sunny Crisp Corn Flakes

N = 100

0

Purchase intent

Viable
95% CI · 61 to 73 Consensus: moderate

Recurring themes

  • + Health cue lands
  • + Shelf standout
  • ! Price hesitation
“I like that it looks less sugary than the big brands. I would try a box if the price was close to what I pay now.”
Simulated shopper No. 072 · F, 34 · Midwest
The Sunny Crisp corn flakes box being tested

Guardrails

The questions we refuse to run.

In an industry that claims 90 percent accuracy, we are upfront about what simulated shoppers can and can't answer. testfeed blocks the questions they can't, before they run and before they cost you anything.

“How much would you pay?”

Stated willingness to pay is unreliable, so we don't ask it.

Rejected

“Rate this 1 to 5.”

Asking an AI for a rating produces inflated scores. Shoppers answer in words only.

Rejected

“How does it taste?”

Simulated shoppers have no senses. Taste, smell and texture are out of scope.

Rejected

“Heard of our brand?”

A simulated shopper has no real memory of your brand, so awareness questions would be meaningless.

Rejected

“What would stop you buying this?”

An open question about purchase barriers. This is the kind that runs well.

Passed
A 1950s focus group: a mix of people and tin robots seated around a low table discussing a plain carton, one robot raising a hand, while a standing moderator takes notes on a red clipboard
  • Unanswerable questions are blocked before they run or cost anything.
  • Results with too few usable responses are suppressed.
  • No winner is declared without a real statistical gap.
  • Every result is labeled as directional, not a market prediction.

Validation

Tested against real human panels.

We have run two head-to-head comparisons against real human panels. Here are both results, including the one we got wrong.

The favorable case · Bae Juice, 2026

Within a few points of a 200-person human panel.

  • Same viable verdict as the human study
  • The same four consumer themes surfaced, independently
  • Gender pattern matched
  • Segment-level accuracy failed, so we don't offer segment-level claims
Viable

The miss · B2B software, 2026

Our B2B software test overshot by 12 points.

  • General-population panel, B2B software stimulus
  • Our internal verdict on our own performance: weak
  • Accuracy depends on the product and the audience being tested
  • This is why every result is labeled directional
Weak

Reproducibility

0%

variation across repeat runs of the same study

Most reliable question

Purchase intent

about three times better aligned with human answers than understanding-style questions

Nine test types

What you can test.

The specifics

Purchase intent

Would they buy it?

Value proposition

Does the pitch land?

Pack test

Which pack wins the shelf?

Price viability

Does the price hold up?

Ad concept

Does the ad persuade?

Brand trust

Do they believe you?

Claims believability

Does the claim convince?

Promotional mechanic

Does the offer tempt?

Event ticket demand

Would they show up?

The methodology

The science behind the verdict.

More detail

No rating scales

Asking an AI to rate something 1 to 5 produces inflated scores. It's a documented failure mode. So shoppers never see a scale. They answer in their own words, one or two sentences at a time.

Semantic Similarity Rating

Each response is embedded and compared against anchor statements for every rating level, producing a probability distribution over ratings. The method is published, independent research from PyMC Labs. The score comes from the language, not from a leading question.

Built on census data

Shoppers start from government census microdata: 100,000 US profiles and 10,000 Australian. Your store's aggregate patterns shape that base. No personal customer data is used.

From answer to score

  1. Words
  2. Embedding
  3. Anchor match
  4. Distribution
  5. Score

Method: Semantic Similarity Rating · published research · PyMC Labs, Maier et al., 2025

Four tin toy robots in cardigans seated at school desks, writing answers in their own words, while a human invigilator holding a red clipboard walks the row

Data and privacy

Your customer data stays private.

Audiences are shaped by aggregate store signals like repeat purchase rates, spend bands, product interests and location. Individual customers are never read or modeled.

  • Read-only Shopify access. testfeed never writes to your store.
  • No pixel and no storefront code. Nothing runs on your live site.
  • Names, emails and addresses are stripped in code before anything is stored.
  • Deletion webhooks purge store data the moment you ask.

White glove service

Custom audiences for bigger brands.

For eight figure brands and FMCG teams, we design bespoke simulated audiences: custom populations built for your category, calibrated against your own data, with studies run for you by our team.

Try it on your own store.

Book a demo and we'll run your first study with you.

Direct install from the Shopify App Store arrives in a few weeks.