Magrios / Knowledge / AI Visibility / Why an all-error run scores null: honest-null be

Why an all-error run scores null: honest-null benchmark design

Guide · AI Visibility · 3 min read · last verified 2026-07-23

Reviewed before publication Editorial board Independent commercial review
In shortZero is measured absence; null is the absence of measurement — and a benchmark that scores failed runs as zero injects fiction into the one trend it exists to protect. The five mechanical rules from Magrios's production tracker, and the…
Why an all-error run scores null: honest-null benchmark design — Magrios diagram
A Magrios diagram — every element states a product fact.

When every call in a benchmark run fails, the run has measured nothing, and its score must be null — not zero. Zero is a finding: the systems were asked, and the brand was absent. Null is the absence of a finding. A benchmark that converts failure into zero manufactures decline that never happened, and once a fake zero enters the series, the trend — the one thing a locked benchmark exists to protect — is no longer evidence.

What is the difference between a zero and a null?

The same digit can be the most honest number in a report or a fabrication, depending entirely on its basis. Magrios's own AI-visibility baseline is a real 0/100, measured twice: providers answered our locked queries, the answers were stored, and the brand simply was not in them. That zero is load-bearing — every future improvement is provable against it.

A null says something categorically different: the instrument did not operate, so the world was not observed. Averaging the two, or displaying them with the same glyph, destroys the distinction that makes either meaningful. The reporting rule follows: a dashboard must be able to show "no data" as a state that is not a number.

How does the rule work mechanically?

Five behaviors, from our production tracker — each one checkable rather than promised:

| Situation | Treatment |

|---|---|

| A provider call fails | An error row is stored, carrying the error message — never a guessed or imputed answer |

| A run has some failures | The aggregate is the mean of scored rows only; error rows are excluded, not zero-filled |

| Every call fails | The run's score is null — no number exists to report |

| A query has only error rows | It is excluded from blind-spot derivation — unmeasured absence is not a finding |

| Mock rows exist for previewing the interface | Stamped as mock in the data and never aggregated with real measurement |

The principle underneath all five: measurement and its failure modes are recorded in the same table, but never blended. The stored error is itself evidence — it says exactly which cells of the run cannot support conclusions. This is the locked benchmark methodology extended to its least glamorous case: locking means the score means one thing, including when that one thing is "nothing was measured."

What goes wrong without honest nulls?

What should a buyer ask a measurement vendor?

Four questions, all answerable with a screen-share rather than a slide:

Frequently asked questions

Why shouldn't a failed benchmark run just score zero?

Because zero asserts measured absence — the systems were asked and the brand was not there — while a failed run observed nothing. Scoring failure as zero manufactures a decline that never happened and a fake recovery when the instrument heals, and once that enters the series, the trend can no longer be trusted as evidence of anything.

How should partial failures within a run be handled?

Store each failure as an error row carrying its error message, exclude those rows from the aggregate, and compute the score as the mean of successfully scored rows only. Never zero-fill: imputed zeros drag the mean by an amount that tracks infrastructure health rather than visibility, so the series quietly starts measuring your uptime instead of your market.

Is a real zero score ever meaningful?

Yes — a measured zero is one of the most useful numbers a young brand can own. Magrios's own baseline is 0/100, measured twice with the run ids recorded: providers answered the locked queries and the brand was absent from every answer. That zero is the fixed point future improvement is proven against, which is exactly why it must never be confused with an unmeasured run.

Further reading — chosen for this article
Entities in this research
Magriosnull resultsbenchmark integritymeasurement errorlocked benchmarktrend corruption
Related knowledge

Measurement questions: what buyers want proven before they pay · linked

The AI visibility gap: how often B2B companies appear in answers about their own market · linked

Benchmark questions: how buyers calibrate what good looks like · linked

Audit-trail requirements for AI-generated recommendations · shared entities

Measuring AI visibility with locked benchmarks · shared entities

Recently updated

Magrios vs Athena · 2026-07-23

Magrios vs Writesonic · 2026-07-23

Magrios vs Semrush · 2026-07-23

Magrios vs peec · 2026-07-23

Where does your brand stand?
Check your AI visibility free — real evidence, not a score.
Check my visibility or run the full analysis →