What is cohort analysis in SaaS
Guide · Glossary & Definitions · 5 min read · last verified 2026-07-29
Cohort analysis is grouping customers by when they started — the month they signed up, the plan they bought, the channel that brought them in — and then following each group separately over time instead of blending everyone into one company-wide number. The payoff is diagnostic: a cohort table can show whether a change actually improved what happens to new customers, something a single blended metric can hide even while it moves in a reassuring direction.
What averages hide
A blended retention number mixes customers who joined two years ago with customers who joined last month, and the mix changes every period as the company adds people. If the product got better in one stretch and worse in another, the two effects can cancel inside the company-wide average — retention looks flat while, underneath, older and newer cohorts move in opposite directions. Averages are not wrong, exactly; they answer a coarser question than the one usually being asked. "Is retention improving" really means "is retention improving for customers who join under current conditions," and only a cohort split answers that version.
The same blending problem shows up whenever the base grows quickly: new customers who have not had time to churn yet dilute downward the risk carried by older customers who have already had every chance to leave. The number can drift for reasons unrelated to product or marketing quality — reason enough to distrust a single trend line and ask what it is made of. Retention and usage are the two metrics this split is built for. Gross Retention vs Net Retention pulls apart the two retention rates a blended figure runs together; What is usage decay? describes the shape a cohort's usage traces as it ages.
How to read a cohort table
A cohort table has one row per starting group and one column per period of age since that group started, not per calendar date. Reading it takes more than one direction.
Read across a row to watch one cohort age:
Cohort Month 0 Month 1 Month 2 Month 3
Jan signups 100% a% b% c%
Feb signups - 100% d% e%
Mar signups - - 100% f%
Read down a column to compare cohorts at the same age — did the March cohort's month-1 figure (d%) beat or trail January's month-1 figure (a%)? That comparison isolates what changed about the product or the buyer, because both groups are measured at an identical point in their life, not at an identical calendar date.
A third read is worth adding once the first two are habit: the diagonal. Every cohort's result in the same calendar month sits on a diagonal line through the table, and a dip that runs along a diagonal rather than down a column usually points to something that hit every customer at once regardless of tenure — an outage, a price change, a broken release — rather than a change in what new customers specifically experience.
What changed in the product vs. what changed in the mix
Once cohorts are separated, two explanations compete for any change in a headline number, and cohort analysis is how they get told apart. The first is a real change in what happens to a customer — onboarding improved, a competitor made switching easier, pricing shifted who qualifies. The second is a change in the mix — the company started selling to a different segment, a stickier channel grew as a share of the total, a partnership brought in a batch that behaves differently from the rest.
Both produce the same movement in a blended number. Only a cohort breakdown, grouped by the dimension suspected of doing the work (signup month, source channel, plan tier, or company size), shows whether the shift lives inside cohorts or between them. If every recent cohort looks like every past cohort at the same age, nothing about the product changed — the company is simply selling to more of one kind of customer than before. If a single cohort's shape changed relative to its predecessors, something happened to the experience itself, on a date that can be found.
The small-cohort trap
A cohort of a handful of customers can look brilliant or disastrous by the accident of two or three accounts, and it will keep looking that way until enough customers pass through it to average out the noise. Treating a single small cohort's early numbers as a verdict — on a new pricing plan, a new segment, a new onboarding flow — is a misuse the shape of the table encourages, because the number arrives looking as authoritative as a large cohort's while the confidence behind it is nothing alike. The guard is patience and volume: wait for enough customers before drawing a conclusion, and when a cohort is unavoidably small, say so next to the number rather than letting the table imply otherwise.
Because a cohort needs time to mature before its later-age columns mean anything, cohort reads are inherently backward-looking — which is the wider case Leading vs Lagging Indicators makes against dashboards built only from mature numbers.
What this is not for
Cohort analysis is descriptive, not causal — it shows that a group behaved differently, not why, and it can look identical whether a product change caused the difference or the cohort simply contained a different kind of customer from the start. Telling those two explanations apart takes an intervention rather than a better table: withhold the change from part of a comparable group and the difference between them is the answer, which is the design a holdout test sets out. A cohort table raises that question; it does not settle it.
A cohort table is also not a defense against the biases that can distort any measurement feeding into it — a table built cleanly from selectively reported or self-graded inputs will still mislead cleanly. Why marketing data flatters itself works through how those inputs get filtered before anyone draws a row. Reading a cohort table correctly is one skill; trusting the data that populated it is a separate, earlier question.
Magrios does not perform this analysis on customer records — it never sees a signup date or an account ID. What it does re-run is the same defined list of buyer questions, checked against AI assistants over time, which produces its own kind of longitudinal read: the same question, asked again later. That produces a trend in what a set of answers cites over time, which is not a cohort of customers grouped and tracked: the material is public answer text and the sources behind it, and no part of it identifies a customer.