Why customer health scores create false confidence
Guide · customer-success · 5 min read · last verified 2026-07-21
Customer health scores create false confidence because they are usually assembled from activity signals — logins, feature usage, support volume, meeting attendance — which measure whether a customer has formed a habit, not whether the customer is receiving value worth renewing.
The gap between habit and value
Activity and value diverge in both directions, and a score built on activity cannot tell the two apart.
A heavily used account can be low value. Teams keep using tools that have become procedural: a weekly export nobody reads, a dashboard checked out of routine, a workflow that survives because removing it requires a decision nobody wants to make. The usage is real. The willingness to pay for it is not.
A lightly used account can be high value. Products consumed episodically — quarterly planning, incident response, annual audit preparation, board reporting — show sparse activity between the moments they matter. The account looks dormant on a score that samples weekly and matters intensely twice a year.
The mechanism behind the false confidence is straightforward. Activity is easy to instrument, so it dominates the inputs. Value is hard to instrument, so it is approximated by activity. The score then reports on its own proxy and is treated as a report on the underlying thing.
How the score misleads
Composite scores fail in specific, repeatable ways.
- Aggregation hides the reason. A score that moves from 80 to 74 does not say which input moved or why. Two accounts at 74 can be in unrelated situations, and the number gives no basis for choosing different responses.
- Weights are set by judgment, then treated as measurement. The relative weight of logins against support tickets is a guess, but the resulting number is presented with the authority of a calculation.
- Green scores suppress investigation. The practical function of a health score is triage: it decides which accounts get attention. A green score removes an account from the queue, and accounts nobody examines are exactly where undetected risk accumulates.
- Offsetting inputs cancel. Rising usage and a departed sponsor can net out to a stable score, and the stability is read as an absence of change when in fact the two most important facts about the account both moved.
- Recency bias distorts. Scores that weight the last few weeks heavily reward a burst of activity in the run-up to renewal, which is often a symptom of a customer preparing an internal review rather than a sign of health.
- The score becomes a target. Once teams are measured on score movement, activity that raises the score gets encouraged regardless of whether it changes anything the customer cares about.
The most consequential failure is the one that is hardest to see from inside the system: a score built on activity will remain green through the entire period in which an account's organizational foundation is dissolving. Sponsor departures, budget consolidation, and reorganizations leave no trace in login data until well after the decision has effectively been made. Those changes drive the churn that arrives without warning.
What activity signals can and cannot support
Activity data is not useless. It is useful for a narrower set of questions than it is usually asked.
Reasonable uses:
- Detecting abrupt discontinuity, such as a workflow that ran daily and stopped entirely.
- Measuring breadth of adoption across teams, which indicates how many people would notice the product's absence.
- Identifying which specific workflows are load-bearing, as an input to renewal conversations.
- Confirming that a newly onboarded customer reached a defined first outcome.
Unreasonable uses:
- Inferring satisfaction. Frequency of use and willingness to pay are different quantities.
- Inferring value. Value is created by what the activity produced downstream, not by the activity itself.
- Forecasting renewal probability on its own.
What a more honest construction looks like
Teams that get more out of health measurement usually change the shape of the artifact rather than tuning its weights.
Keep signals separate and visible. Replacing one composite number with a small set of named facts — sponsor status, number of dependent workflows, most recent delivered outcome, commercial posture — preserves the reason behind the assessment. Each fact points to a different response.
Weight organizational facts above behavioral ones. Who owns the budget, who would defend the renewal internally, and how many teams depend on the product are more decision-relevant than session counts.
Measure outcomes produced, not actions taken. Where the product creates an artifact the customer uses elsewhere — a report that enters a planning cycle, an alert that changes a decision — the existence and consumption of that artifact is a better indicator than the clicks that generated it.
Track dependency, not frequency. The renewal question is what breaks if the product disappears. An account with three deeply embedded workflows and modest usage is more defensible than an account with broad shallow usage across many features, and this distinction is invisible in a frequency-weighted score.
Validate against outcomes. Compare scores at a fixed point before renewal against what actually happened. A score that is never checked this way carries weights that persist for years without evidence.
Record confidence explicitly. An account nobody has spoken with in months should not carry the same certainty as one with a recent substantive conversation. Marking staleness prevents old information from being read as current assessment.
Where health measurement still helps
The deeper correction is to stop treating the score as an answer. Health measurement works as a mechanism for deciding where to look, and fails as a substitute for having looked. When a green score is accepted as evidence that an account is fine, the score has replaced the investigation it was built to prioritize — and the accounts that eventually surprise the forecast are the ones nobody was examining: the ones that stayed green. Accounts scored on activity alone will produce a retention forecast that looks stable right up to the point where several renewals resolve differently than expected, with visible consequences for net revenue retention and for the lifetime value assumptions built on top of it.