Why an all-error run scores null: honest-null benchmark design
Guide · AI Visibility · 3 min read · last verified 2026-07-23

When every call in a benchmark run fails, the run has measured nothing, and its score must be null — not zero. Zero is a finding: the systems were asked, and the brand was absent. Null is the absence of a finding. A benchmark that converts failure into zero manufactures decline that never happened, and once a fake zero enters the series, the trend — the one thing a locked benchmark exists to protect — is no longer evidence.
What is the difference between a zero and a null?
The same digit can be the most honest number in a report or a fabrication, depending entirely on its basis. Magrios's own AI-visibility baseline is a real 0/100, measured twice: providers answered our locked queries, the answers were stored, and the brand simply was not in them. That zero is load-bearing — every future improvement is provable against it.
A null says something categorically different: the instrument did not operate, so the world was not observed. Averaging the two, or displaying them with the same glyph, destroys the distinction that makes either meaningful. The reporting rule follows: a dashboard must be able to show "no data" as a state that is not a number.
How does the rule work mechanically?
Five behaviors, from our production tracker — each one checkable rather than promised:
| Situation | Treatment |
|---|---|
| A provider call fails | An error row is stored, carrying the error message — never a guessed or imputed answer |
| A run has some failures | The aggregate is the mean of scored rows only; error rows are excluded, not zero-filled |
| Every call fails | The run's score is null — no number exists to report |
| A query has only error rows | It is excluded from blind-spot derivation — unmeasured absence is not a finding |
| Mock rows exist for previewing the interface | Stamped as mock in the data and never aggregated with real measurement |
The principle underneath all five: measurement and its failure modes are recorded in the same table, but never blended. The stored error is itself evidence — it says exactly which cells of the run cannot support conclusions. This is the locked benchmark methodology extended to its least glamorous case: locking means the score means one thing, including when that one thing is "nothing was measured."
What goes wrong without honest nulls?
- Phantom decline. A provider outage becomes a scored collapse; the next healthy run becomes a recovery. Both movements are fiction, and both trigger real work — postmortems for a decline that never occurred, credit for a rebound that was just the instrument healing. This is sampling error's nastier cousin: not noise around a true value, but signal invented from no observation at all.
- Silent mean-dragging. Zero-filling partial failures pulls every mixed run downward by an amount that tracks infrastructure health, not visibility. The series stops measuring the market and starts measuring your uptime.
- Alert fatigue with a memory. Teams learn that drops might be outages, so real drops get triaged skeptically. The benchmark's authority — the reason to lock it in the first place — erodes from the inside.
- Corrupted comparisons. A measurement window containing a fake zero cannot be honestly compared with any other window, and nothing downstream can repair that, because the corruption is indistinguishable from data.
What should a buyer ask a measurement vendor?
Four questions, all answerable with a screen-share rather than a slide:
- How are provider errors scored? The only acceptable answer is that they are not — errors are excluded and shown as errors.
- What does a fully failed run display? If the answer is any number, including zero, the trend contains fiction.
- Is demo or mock data distinguishable from measurement in the stored data itself — not just visually, so it can never be aggregated by accident?
- Can your interface show the difference between "we measured absence" and "we did not measure"? A tool that cannot render that distinction cannot report your worst week honestly — and the vendor's willingness to face its own bad numbers is the fastest available proxy. Ours is 0/100, measured twice, and we publish the method that produced it.