Launching soon on the Shopify App Store

testfeed

concept-testing

Concept Testing: The Definitive Guide

Concept Testing: The Definitive Guide

Most concept tests fail for a reason nobody admits: the test could not tell a good idea from a bad one. The concept came back with a decent score, the team felt reassured, the product launched, and it sank anyway. The problem was never the idea. It was a test designed to produce comfort rather than a decision.

Concept testing is the market research method that shows an idea to your target buyers before you build it or spend on it, and measures whether they would actually part with money. Done well, it is the cheapest insurance you can buy against a bad launch. Done badly, it is a very expensive way to confirm what you already believed. This guide is about the difference.

What concept testing actually is

Concept testing puts an early-stage idea in front of the people you want to sell to, before full development, and measures how well it lands. The idea can be almost anything you would otherwise commit money to on instinct: a new product, a pack design, an ad, a positioning claim, a price, a product name, a promotion, or the demand for an event. You present the concept, you ask a structured set of questions, and you come away with a read on whether it is worth backing. When the idea is a physical product, people often call this product concept testing, but the logic is identical whatever the format.

The word people forget is “decision”. A concept test is not a research exercise you tick off. It exists to answer one question: do we develop this, refine it, or kill it? Everything about how you design the test should serve that decision, and if a test cannot change what you do next, it is theatre.

It helps to clear up a myth that gets used to justify all this. You have probably heard that 80 to 95 per cent of new products fail, so you had better test everything. That figure is closer to folklore than fact. When George Castellion and Stephen Markham reviewed the actual evidence across nineteen peer-reviewed studies for the Journal of Product Innovation Management, they found real new-product failure rates sit around 40 per cent, not 80. The scary number survives because it is repeated by people with an interest in selling you research and consulting.

So the honest case for concept testing is not fear. It is allocation. Roughly two in five new products will not work, you cannot tell in advance which two, and your budget only stretches to a few real bets. Concept testing tips the odds by cutting the weak ideas before they cost you a development cycle and an ad budget. It does not guarantee a winner. It improves the ratio.

When to run a concept test, and when not to

Concept testing belongs at the front end of product development. In Robert Cooper’s Stage-Gate model, the idea-to-launch framework most consumer businesses still run on, the early gates exist precisely to kill weak concepts before they consume development resource. That is the work concept testing does: it feeds the go, kill or hold decision at a gate, when changing course is still cheap.

In practice there are four moments where it earns its place. The first is screening, when you have more ideas than budget and need to rank them. The second is validation, when one idea is the frontrunner and you want to pressure-test it before committing to development. The third is refinement, when a concept is promising but you are choosing between versions, such as three routes for a pack or two price points. The fourth is pre-spend, when the product exists but the ad, the claim or the launch angle does not, and you want to know which creative direction to fund.

There are also times to skip it. If you will launch regardless of the result, do not run the test; you are only buying yourself a number to ignore. If the decision rests on taste, texture or smell, a concept test cannot help you, because no amount of describing a flavour tells you whether it is nice. And if you have no real audience to ask, testing on your team or your friends is worse than not testing, for reasons I will come to. Concept testing answers “would people want this”. It does not answer “is this well made” or “will my mate be polite about it”.

The concept testing methods, and how to choose

There are four standard concept testing methods. The differences look academic until you realise each one changes the answer you get.

Monadic testing shows each respondent a single concept in isolation. They never see the alternatives. Because nobody is comparing, the score you get reflects how the concept performs on its own merits, which is how it will meet the market: a shopper in the aisle is not holding your three rejected pack designs. Monadic is the cleanest, most realistic read. The cost is sample. You need a fresh group of people for every concept, so testing five ideas means five times the respondents.

Sequential monadic testing shows each respondent several concepts, one after another, in a rotated order. It is far cheaper, because the same people evaluate everything, and rotation controls for the fact that whatever comes first tends to score differently from whatever comes last. The trade-off is contamination: once someone has seen concept A, they cannot un-see it when they judge concept B. Fatigue creeps in after the first two or three. For screening a longer list it is the pragmatic choice, as long as you keep the list short and rotate properly.

Comparative testing puts the concepts side by side and asks people to choose. It is intuitive and it feels decisive, which is exactly why it is dangerous. Direct comparison exaggerates small differences and creates preferences that would never surface in real life, where your concept competes against the category, not against your other drafts. Comparative testing answers “which of these do you prefer”, which is rarely the question that matters. Use it for a genuine head-to-head, such as two final name candidates, and distrust the margins.

Protomonadic testing runs a monadic read first, then asks the comparison at the end. You get the realistic in-isolation score and the direct preference, and you can check whether they agree. When the monadic winner and the comparative winner are the same concept, you can move with confidence. When they disagree, you have learned something important about how context changes the choice. It is the most reliable design and the most expensive, so save it for high-stakes decisions.

Here is the short version for when you are deciding which to run.

MethodBest forSample costMain risk
MonadicA realistic read on one or a few conceptsHigh (fresh sample per concept)Cost at scale
Sequential monadicScreening a short list on a budgetLow (one sample sees all)Order and fatigue effects
ComparativeA true final head-to-headLowOverstates differences vs real life
ProtomonadicHigh-stakes launch decisionsHighestComplexity, cost

If you take one thing from this section: default to monadic, reach for sequential monadic when budget forces your hand, and treat comparative results as directional at best.

What to measure: the metrics that actually predict

A concept test lives or dies on what you ask. Score the wrong things and you get a confident answer to a question that does not matter. Five or six dimensions do most of the work.

Purchase intent is the headline. It is the closest proxy you have for the behaviour you care about, and it is the number most tied to the eventual decision. Ask how likely someone is to buy, on a scale, and report the top-two-box: the share of people choosing the top two points, for example “definitely would buy” plus “probably would buy”. Top-two-box is more stable than an average and easier to benchmark.

Appeal captures whether people simply like it. Uniqueness or distinctiveness tells you whether it stands out from what is already on the shelf, which matters enormously for cut-through. Believability checks whether the promise is credible, because a claim people do not believe will not convert however appealing it sounds. Relevance asks whether it meets a real need for this person. And value for money tests whether the price and the perceived worth line up.

The pattern to watch for is the concept that scores well on appeal but poorly on uniqueness. People like it, but it does not stand out, so it will get lost. The winners tend to score high on both purchase intent and uniqueness at the same time. That is the quadrant worth backing: wanted and different. High intent with low uniqueness is a me-too. High uniqueness with low intent is a curiosity nobody will buy.

Now the part most guides skip, because it is inconvenient. Stated purchase intent overstates real behaviour, systematically and by a lot. People say they will buy far more often than they do. The academic record on this is old and consistent. In a classic experiment on survey design, economist F. Thomas Juster showed that purchase-probability scales explain roughly twice as much of the variance in actual buying as simple yes-or-no intention questions, because most purchases come from people who never declared an intention at all. The broader meta-analysis of when intentions predict sales, by Vicki Morwitz, Joel Steckel and Alok Gupta in the International Journal of Forecasting, found the link is weaker for new products than for established ones, and weaker the longer the gap between the survey and the shelf. Concept testing sits in the worst quadrant for this: new products, measured well before purchase.

The practical takeaway is not “ignore purchase intent”. It is “never read it as a sales forecast”. Treat the top-box number as a relative signal, compare it against your own history or a known benchmark concept rather than an absolute pass mark, and where you can, ask people to estimate a probability rather than tick “definitely”. You are ranking ideas against each other, not predicting units.

If you must translate intent into a rough demand read, apply a reality discount rather than taking the number at face value. A common working rule is to keep only about half to two-thirds of top-two-box intent as a real-world estimate, and to discount “definitely would buy” far less harshly than “probably would buy”, since the second box is where most of the overstatement hides. The exact multiplier should come from your own history once you have a few tests to calibrate against. Until then, assume the number is generous.

How to write the concept, the part everyone rushes

Garbage in, garbage out applies with brutal force here. A concept test can only be as good as the concept statement you put in front of people, and a leading, salesy write-up will hand you a flattering, useless result.

A clean concept statement has a simple structure. Open with the insight or the problem, the reason this should exist. State the idea plainly. Give the main benefit, the thing it does for the person. Add the reason to believe, the proof that the benefit is real. And include the price, because a concept without a price tests a fantasy. Keep the language neutral. The moment you write “revolutionary” or “the only product that”, you are no longer testing the idea, you are testing your own marketing, and everyone rates marketing more harshly than they rate a plain description.

Test the concept, not the execution. At this stage you want to know whether the idea has legs, not whether this particular headline or that shade of blue works. Strip it back to the proposition. You can test the execution later, once you know the concept itself is worth executing.

Set your action standard before you run the test

This is the single step that separates a concept test that works from one that is, to borrow the phrase from a well-known thread of frustrated researchers, worthless. Decide your pass or fail threshold before you see a single response.

Write it down. Something like: a concept goes forward if top-two-box purchase intent clears our benchmark and uniqueness is above the midpoint; otherwise we refine or cut. The specific numbers depend on your category and your own norms, and that is fine. What matters is that the rule exists before the data does. The reason is human, not statistical. Once the results are in, everyone in the room can construct a story for why their favourite concept deserves a pass despite the score. An action standard set in advance is the only thing that stops a concept test collapsing into a debate you were always going to win for the idea you already liked.

This is also what makes the test honest about limits. If nothing clears the bar, that is a result, not a failure of the research. It means you have not found the idea yet, and you have learned it for the price of a survey rather than a launch.

A concept testing example, worked through

Take a direct-to-consumer snack brand choosing between three flavour concepts for a new line. Call them Concept A, B and C. The team runs a monadic test, a fresh sample of around 150 category buyers per concept, and sets the action standard in advance: go forward only if top-two-box purchase intent beats the brand’s benchmark of 45 per cent and uniqueness scores above the scale midpoint. The numbers below are illustrative, to show the logic rather than to quote a real study.

Metric (top-two-box)Concept AConcept BConcept C
Purchase intent52%61%44%
UniquenessHighLowHigh
Value for money48%55%40%

At a glance, Concept B looks like the winner. It has the highest purchase intent and the best value read. A team without an action standard would ship B and feel good about it. But apply the rule. B clears the intent bar comfortably, yet it scores low on uniqueness: people like it because it is familiar, and it will disappear on a shelf next to three established brands doing the same thing. Concept A clears both bars, wanted and distinctive, which is the quadrant that survives contact with the market. Concept C fails the intent threshold and is out.

The disciplined call is to develop A, hold B as a possible line extension where distinctiveness matters less, and drop C. Notice what the method saved you from. Had you run a comparative test instead, B’s familiarity would likely have won the head-to-head and buried A, because comparison rewards the safe option. The monadic design, read against a pre-set standard, pointed you at the idea with a real chance of standing out. That is concept testing doing its job.

Sample size and who you actually ask

Two questions come up constantly: how many people, and which people. The second matters more than the first.

On numbers, a common working benchmark for a directional read is 100 to 200 completed responses per concept. For a high-stakes launch you want more, 300 or more per concept, with demographic quotas so the sample mirrors your real buyers. Below about 100 per cell you are reading noise.

On who, the rule is simple and routinely ignored: ask category buyers, not warm bodies. A hundred and twenty people who actually buy in your category will tell you more than five hundred random respondents, and infinitely more than your team, your investors or your friends, all of whom are compromised by knowing you. Screen for the behaviour that matters, buying the category in the last few months, and be ruthless about it. The most expensive mistake in concept testing is not too small a sample. It is a large sample of the wrong people. For a fuller treatment of building and testing against the right audience, the TestFeed guide to audience testing goes deeper.

The mistakes that make concept testing worthless

Most failed concept tests share the same handful of errors, and they are all avoidable.

Leading concept copy is the first, writing the statement like an ad so the idea flatters itself into a good score. No action standard is the second, running the test without a pre-set threshold so the result becomes whatever the loudest person in the room wants it to be. The wrong audience is the third, testing on people who are not real buyers. Reaching for comparative testing when the decision needs a realistic in-isolation read is the fourth, and it quietly inflates differences that will not exist in the market. Treating top-box purchase intent as a sales forecast is the fifth, and it sets you up to be disappointed by numbers that were only ever relative. And testing what you cannot judge from a concept is the sixth: no survey settles whether a snack tastes good.

Avoid those six and you are already ahead of most teams running concept tests, including well-funded ones.

Faster and cheaper ways to test concepts now

The old model of concept testing meant a research agency, a bespoke study and a four-to-six-week wait, which priced out most growing brands and slowed down the ones who could afford it. That has changed. Concept testing tools now fall into a few groups: traditional panel providers with normative databases, self-serve survey platforms with consumer panels, AI-moderated qualitative tools, and AI or synthetic-audience platforms. Each carries a different trade-off on cost, speed and depth, which I have compared in detail in the rundown of concept testing platforms.

There is a second-order effect here that most guides miss. The whole monadic-versus-comparative trade-off above rests on sample cost: monadic gives the cleanest read, but you pay for a fresh sample per concept, which is what pushes budget-constrained teams toward the cheaper, more biased sequential and comparative designs. When a synthetic or AI audience makes a fresh sample per concept effectively free, that calculation changes. You can run pure monadic across all five ideas for roughly the cost of one, and stop trading realism for budget at the screening stage. The bias you were accepting to save money was never worth it; it is just that, until recently, avoiding it was expensive.

The shift that matters for a founder or a small insights team is speed and cost at the screening stage, where you are killing weak ideas and do not yet need agency-grade rigour. This is where an AI approach fits. TestFeed lets you test a product, a pack, an ad, a claim, an in-context price, a concept, a name or a promo against an audience modelled on your buyers, and get back a purchase-intent read, a plain verdict, the shoppers’ reasons in their own words, and a clear next move, in days rather than weeks. The point is a fast, directional signal at the front end, so you can cut the losers and take only the strongest concepts into a fuller study.

Be clear-eyed about what that is and is not. It is pre-launch, pre-spend signal, not a market forecast or a guaranteed sales number, and it does not judge taste, texture or smell. It is a first pass that earns its place by being fast and cheap enough to run on every idea, so the weak ones never reach an expensive study. For a final, high-stakes launch decision, pair it with a recruited panel. Used that way, the faster tools do not replace rigour. They stop you wasting rigour on ideas that were never going to make it. If you want the wider view of where AI fits in research before you commit, the guide to AI market research covers the full picture, and the product launch strategy guide shows where concept testing sits in the run-up to launch.

Frequently asked questions

What is concept testing?

Concept testing is a research method that puts an early-stage idea, such as a product, pack, ad, name or price, in front of your target audience before full development, and measures how well it resonates. The goal is a clear decision: which concepts to develop, which to refine, and which to cut before you spend on production or launch.

What are the main concept testing methods?

The four standard methods are monadic (each person sees one concept in isolation), sequential monadic (each person sees several concepts in rotated order), comparative (concepts shown side by side), and protomonadic (a monadic read followed by a direct comparison). Monadic gives the cleanest, most realistic read. Sequential monadic is cheaper. Comparative is the most prone to overstating differences.

How many respondents do I need for a concept test?

For a directional read, 100 to 200 completed responses per concept is a common working benchmark. High-stakes launch decisions typically use 300 or more per concept with demographic quotas. The audience matters more than the number: 120 real category buyers beat 500 random people every time.

What is a good purchase intent score in concept testing?

There is no universal pass mark, because scores vary by category, scale and audience. What matters is the top-two-box figure, the share choosing the top two points on the intent scale, measured against your own norms or a benchmark concept, and read as a directional signal rather than a sales forecast. Stated intent consistently overstates real purchasing, so discount the top box rather than taking it at face value.

What is the difference between concept testing and usability testing?

Concept testing asks whether an idea is worth building, before it exists. Usability testing asks whether a built product or interface is easy to use. Concept testing happens at the front end of development; usability testing happens close to or after launch. Different questions, different methods, different stage.

Where to start

If you run one concept test this quarter, get two things right before you worry about anything else. Ask real category buyers, not the people around you. And write your action standard down before you look at the data. The methods, the sample sizes and the tools all matter, but those two habits are what separate a concept test that changes your decision from one that just makes you feel better about the decision you had already made. Pick the idea you are least sure of, write the rule, and test it this week.

Millie Marconi

Written by

Millie Marconi

CEO & Co-Founder, TestFeed

Millie is a market researcher and former ecommerce store owner who has worn just about every hat in marketing. She writes about AI, customer research and ecommerce.

Try it on your own store.

Book a demo and we'll run your first study with you.

Direct install from the Shopify App Store arrives in a few weeks.