How to document growth experiments
Guide · Frameworks · 4 min read · last verified 2026-07-27
An experiment record is a dated, written account of a growth bet: what the team believed, what evidence motivated the belief, what result would have counted as disconfirmation, what actually happened, and what was decided. The definition matters because most teams think they document experiments when what they actually keep is a task list with outcomes — and a task list cannot stop the same failed idea from being proposed again, eighteen months later, by someone who was not there the first time.
Why teams repeat failed experiments
Learnings default to living in people, and people leave, change roles, and forget. The pattern is predictable: an experiment fails, the lesson lodges in one operator's memory, that operator departs, and the next hire — reasoning correctly from the same visible surface — proposes it again. Nothing in the environment records that the ground was already tried and found barren. Failure is also socially under-remembered: teams retell their wins and quietly stop mentioning the campaigns that fizzled, so the organization's oral history skews toward whatever flattered it. A written archive is the only durable correction. The same argument holds one level up — a company that re-buys the same market research every time context walks out the door has the identical disease, which is why intelligence about a market needs a memory for exactly the reasons experiments do.
The five fields
A usable record has five fields, and their order is the discipline.
Hypothesis. One falsifiable sentence: we believe this change will produce this effect for this segment. If the sentence cannot fail, it is a hope rather than a hypothesis, and the record will be unreadable later because nobody can say what was tested.
Motivating evidence. Why this bet, and why now — the observation that triggered it: a pattern in buyer questions, a gap in how assistants answer a category question, a lost-deal reason that keeps recurring. This field lets a future reader judge whether the conditions that justified the bet still hold before rerunning it.
What would change our mind. The pre-committed disconfirmation condition, written before launch: an observable state of the world, with a date attached, that the team agrees in advance counts as the hypothesis failing.
Result. What was observed, stated in the same terms the previous field used, so the comparison is mechanical rather than rhetorical.
Decision. Kill, keep, scale, or rerun with a named change — plus one sentence of why. A record that ends without a decision is a story, and stories accumulate without compounding.
Write the falsifier before the result
The third field is the load-bearing one, and its timing is the entire point. Written after the fact, success conditions migrate toward whatever happened — hindsight is a fluent editor, and it edits honest people too. Written before, the falsifier makes the result field mechanical: compare the observation against the pre-commitment and the verdict largely writes itself, no rhetorical skill required. This is the same measurement discipline as locking a benchmark before acting on it — the logic worked through in this guide to measuring answer change — applied at the scale of a single bet. A team adopting only one habit from this page should adopt this one.
A template you can copy
The record should take minutes to open and minutes to close. A plain-text form is enough:
Experiment: <short name> — opened <date> — owner <name>
Hypothesis: We believe <change> will <effect> for <segment>.
Motivating evidence: <the observation that triggered this, with a link>
We will change our mind if: <observable condition> by <date>.
Result (<date>): <what was observed, in the falsifier's terms>
Decision: kill / keep / scale / rerun — <one sentence of why>
The angle brackets are left empty of invented content on purpose. The value of a record is entirely in being filled with your actual evidence, and a template pre-filled with plausible-looking specifics teaches teams to write plausible-looking records.
Where the archive lives and when it gets written
One searchable place, adjacent to the evidence it cites — not scattered across chat threads, slide appendices, and personal notes, where search dies and departures take the index with them. The writing has a natural cadence: the first three fields are completed the day the experiment opens, and the last two are completed the day it receives its verdict — which, in a healthy operating rhythm, happens inside the weekly growth review, where one running experiment is decided each week. At quarter's end the archive becomes the primary input to planning the next 90-day loop: the kills mark ground already covered, the keeps mark motions worth funding harder, and the open records mark commitments the new quarter inherits rather than forgets.
The archive pays out on read
Documentation disciplines die when they are write-only. The payoff moments are reads. Before proposing an experiment, search the archive for prior attempts on the same ground. When onboarding a growth hire, hand over the archive as the fastest honest history of what the company learned the hard way. When a strategy debate stalls on competing recollections, let a dated record settle it. Institutional memory is not a filing habit; it is the ability to make this quarter's decisions with last year's evidence in the room. Magrios applies the same principle to market position — persistent question sets and locked benchmarks so that change is measurable rather than remembered — and the experiment archive is the team-sized version of that commitment. It costs a few sentences per bet and repays them the first time someone almost repeats a failure.