What is a holdout test in marketing
Guide · Glossary & Definitions · 5 min read · last verified 2026-07-29
A holdout test is deliberately withholding a marketing activity — a campaign, an offer, a channel — from a comparable slice of the audience, so the excluded group's outcome shows what would likely have happened anyway. The gap between the two groups, treated and held out, is about as close as a marketing team can get to knowing whether an activity caused a result rather than merely preceding one. It is also, by design, expensive: knowing requires not doing something to a group large enough for the comparison to mean anything.
What makes a holdout valid
Three conditions have to hold or the comparison stops measuring what it claims to. The two groups have to be comparable before the test starts — assigned by something close to random, not by a rule that happens to correlate with the outcome, like holding out only customers who did not respond to the last few campaigns. There has to be enough volume in both groups for a real difference to be visible over whatever noise the market produces on its own in a given week. And the test needs patience: cutting it short the moment the treatment group pulls ahead invites reading normal week-to-week variation as a settled result.
Miss the comparability condition and the test answers a different question than intended — a holdout built from an already-different population measures the difference between two kinds of customers, not the effect of the campaign. Miss volume or patience and the test technically ran but produced a number too noisy to trust either way.
What you are paying for
Running a holdout costs the outcome that group would otherwise have produced. If the campaign works, the held-out group represents revenue not captured, on purpose, for the length of the test. That is not a flaw in the method — it is the method; you cannot see the world where nothing was done without actually doing nothing to part of it. The cost is real and specific enough to budget for, which is also what makes it easy to leave unnamed: spending visible reach to answer a question the team already has a comfortable guess about is an uncomfortable trade to propose out loud.
The size of the holdout sets the size of the bill. A small excluded group costs little and answers little, because the comparison stays noisy; a larger one answers more clearly and costs more of the outcome it is measuring. No volume makes the test free — only a volume that makes the answer worth what it costs.
What the bill buys also depends on who is in the excluded group, not just how many. Withholding a campaign from a slice of a high-intent audience costs more than withholding it from a cold one of the same size, which makes the cheap holdout tempting: draw the excluded group from the part of the audience with the least to lose. That shortcut breaks the first condition by construction — a group selected for having less at stake is no longer comparable to the group that kept the campaign, and the test now measures the selection rather than the activity.
When "we cannot afford to know" is the right call
Sometimes the right answer to did-this-work is that finding out costs more than the campaign itself. A small channel, a short-lived promotion, or a market too thin to support two comparable groups are all cases where a holdout would consume a meaningful share of the very audience the activity is trying to reach, for a result too noisy to trust once measured. In those cases, running the test is worse than not knowing — it spends real reach to produce an answer with an error band wide enough to be nearly useless.
The trade is worth making explicitly rather than by default. Before ruling a holdout out, name what a wrong guess would cost if the activity does not work and continues unchecked; only when that ongoing cost is smaller than the cost of testing does skipping the test hold up as the better trade, rather than simply the easier one.
Holdout vs. the neighboring methods
A holdout is one way to answer a question that has a name of its own — incrementality — and it helps not to confuse the method with the concept it serves. Incrementality is the question (would this have happened anyway); a holdout is one design for answering it, alongside geographic splits and staged rollouts that build the comparison group differently.
Attribution is a different tool entirely, running continuously rather than episodically. Attribution modeling is where the crediting rules themselves are set out; how an occasional test calibrates those rules is already argued in the incrementality piece above, and saying it a second time here would add nothing to it.
Cohort analysis is different again — it tracks a group's behavior over time without withholding anything from anyone, useful for spotting that something changed but unable, alone, to say a marketing activity caused the change the way a controlled holdout can. For how to read one, start with What is cohort analysis.
Where this can still go wrong
A holdout is not immune to the failures that distort other measurements — a split that leaks, where someone in the holdout group sees the campaign anyway through a shared inbox or a forwarded link, or a result reported only when it flatters the channel being tested, can make a holdout mislead as convincingly as any other number. The same corrective instinct applies here as anywhere else data gets reported: ask what would have to be true for this comparison to be clean. That habit generalizes; Why marketing data flatters itself applies it to numbers that never went near a test. A result worth trusting also deserves a more durable home than a slide, which is the case How to document growth experiments makes for recording the design, the dates, and the conditions that held while it ran.
Magrios does not run holdout tests. It has no mechanism for withholding anything from a comparable group, because it is not a campaign tool — it is a repeated scan of how AI assistants answer real buyer questions about a company. What it can show is whether a citation picture changed between two scans; whether that change caused a downstream business result, rather than merely accompanying one, is the kind of question a holdout is built to separate out.