Magrios / Knowledge / Glossary & Definitions / What is a holdout test in marketing

What is a holdout test in marketing

Guide · Glossary & Definitions · 5 min read · last verified 2026-07-29

Reviewed before publication Editorial board Independent commercial review
In shortA holdout test is deliberately withholding a marketing activity — a campaign, an offer, a channel — from a comparable slice of the audience, so the excluded group's outcome shows what would likely have happened anyway.

A holdout test is deliberately withholding a marketing activity — a campaign, an offer, a channel — from a comparable slice of the audience, so the excluded group's outcome shows what would likely have happened anyway. The gap between the two groups, treated and held out, is about as close as a marketing team can get to knowing whether an activity caused a result rather than merely preceding one. It is also, by design, expensive: knowing requires not doing something to a group large enough for the comparison to mean anything.

What makes a holdout valid

Three conditions have to hold or the comparison stops measuring what it claims to. The two groups have to be comparable before the test starts — assigned by something close to random, not by a rule that happens to correlate with the outcome, like holding out only customers who did not respond to the last few campaigns. There has to be enough volume in both groups for a real difference to be visible over whatever noise the market produces on its own in a given week. And the test needs patience: cutting it short the moment the treatment group pulls ahead invites reading normal week-to-week variation as a settled result.

Miss the comparability condition and the test answers a different question than intended — a holdout built from an already-different population measures the difference between two kinds of customers, not the effect of the campaign. Miss volume or patience and the test technically ran but produced a number too noisy to trust either way.

What you are paying for

Running a holdout costs the outcome that group would otherwise have produced. If the campaign works, the held-out group represents revenue not captured, on purpose, for the length of the test. That is not a flaw in the method — it is the method; you cannot see the world where nothing was done without actually doing nothing to part of it. The cost is real and specific enough to budget for, which is also what makes it easy to leave unnamed: spending visible reach to answer a question the team already has a comfortable guess about is an uncomfortable trade to propose out loud.

The size of the holdout sets the size of the bill. A small excluded group costs little and answers little, because the comparison stays noisy; a larger one answers more clearly and costs more of the outcome it is measuring. No volume makes the test free — only a volume that makes the answer worth what it costs.

What the bill buys also depends on who is in the excluded group, not just how many. Withholding a campaign from a slice of a high-intent audience costs more than withholding it from a cold one of the same size, which makes the cheap holdout tempting: draw the excluded group from the part of the audience with the least to lose. That shortcut breaks the first condition by construction — a group selected for having less at stake is no longer comparable to the group that kept the campaign, and the test now measures the selection rather than the activity.

When "we cannot afford to know" is the right call

Sometimes the right answer to did-this-work is that finding out costs more than the campaign itself. A small channel, a short-lived promotion, or a market too thin to support two comparable groups are all cases where a holdout would consume a meaningful share of the very audience the activity is trying to reach, for a result too noisy to trust once measured. In those cases, running the test is worse than not knowing — it spends real reach to produce an answer with an error band wide enough to be nearly useless.

The trade is worth making explicitly rather than by default. Before ruling a holdout out, name what a wrong guess would cost if the activity does not work and continues unchecked; only when that ongoing cost is smaller than the cost of testing does skipping the test hold up as the better trade, rather than simply the easier one.

Holdout vs. the neighboring methods

A holdout is one way to answer a question that has a name of its own — incrementality — and it helps not to confuse the method with the concept it serves. Incrementality is the question (would this have happened anyway); a holdout is one design for answering it, alongside geographic splits and staged rollouts that build the comparison group differently.

Attribution is a different tool entirely, running continuously rather than episodically. Attribution modeling is where the crediting rules themselves are set out; how an occasional test calibrates those rules is already argued in the incrementality piece above, and saying it a second time here would add nothing to it.

Cohort analysis is different again — it tracks a group's behavior over time without withholding anything from anyone, useful for spotting that something changed but unable, alone, to say a marketing activity caused the change the way a controlled holdout can. For how to read one, start with What is cohort analysis.

Where this can still go wrong

A holdout is not immune to the failures that distort other measurements — a split that leaks, where someone in the holdout group sees the campaign anyway through a shared inbox or a forwarded link, or a result reported only when it flatters the channel being tested, can make a holdout mislead as convincingly as any other number. The same corrective instinct applies here as anywhere else data gets reported: ask what would have to be true for this comparison to be clean. That habit generalizes; Why marketing data flatters itself applies it to numbers that never went near a test. A result worth trusting also deserves a more durable home than a slide, which is the case How to document growth experiments makes for recording the design, the dates, and the conditions that held while it ran.

Magrios does not run holdout tests. It has no mechanism for withholding anything from a comparable group, because it is not a campaign tool — it is a repeated scan of how AI assistants answer real buyer questions about a company. What it can show is whether a citation picture changed between two scans; whether that change caused a downstream business result, rather than merely accompanying one, is the kind of question a holdout is built to separate out.

Frequently asked questions

What is a holdout test?

Ad platforms use terms like 'ghost ad' and 'PSA test' for closely related designs, and inside a growth team the same idea may simply be called a control group. They are not one structure: they differ in how the control group gets built. A PSA test shows the held-back group a public-service ad in place of the real one, substituting rather than withholding, so whatever the substitute does to attention rides along in the result. A ghost ad records which control users the system would have served, matching on eligibility instead of excluding at random — a different answer to the comparability problem, not the same one. A plain holdout withholds and does nothing else. One practical limit is worth naming: holdouts work best for activities customers do not expect and would not miss, like a discre

How do holdout groups prove marketing works?

Strictly, a holdout does not produce proof — it produces a difference between two groups, plus a sense of how likely that difference is to be real rather than noise, which is a range, not a yes-or-no verdict. A single test nudges belief in one direction; repeated tests, run cleanly across different periods, narrow the range rather than close it. The estimate stays an estimate: business-to-business volumes are small and market noise is not, so even a clean split leaves a band rather than a point. Treating one holdout result as final proof is an overreach the method does not support.

When is a holdout worth the cost?

Weigh it against the size of the decision riding on the answer, not the size of the campaign. A holdout is worth running before a decision that is expensive to reverse — renewing an annual channel contract, doubling a budget line, killing a channel part of the team still believes in — because the cost of the test is small next to the cost of committing further on a wrong guess. For a routine campaign that can simply be adjusted next month if it underperforms, the ongoing result usually teaches enough without the added cost of a formal holdout.

Further reading — chosen for this article
Entities in this research
Magriosholdout testincrementalityexperimentsdefinition
Related knowledge

What is a customer data platform · shared entities

What is a creative brief in marketing · shared entities

What is self-reported attribution · shared entities

Recently updated

Where to expand internationally first · 2026-07-29

What is a UTM parameter · 2026-07-29

What is a right-to-audit clause · 2026-07-29

What is an order form in SaaS deals · 2026-07-29

Where does your brand stand?
Check your AI visibility free — real evidence, not a score.
Check my visibility or run the full analysis →