Magrios / Knowledge / SEO / AEO / GEO / Measuring AI visibility with locked benchmarks

Measuring AI visibility with locked benchmarks

Guide · SEO / AEO / GEO · 3 min read · last verified 2026-07-22

Reviewed before publication Editorial board Independent commercial review
In shortComparable AI-visibility measurement requires locking the queries and the prompt, storing answers verbatim, and refusing fake scores. The working design of Magrios's own tracker — the one that scored us 0/100 twice — both runs kept on the…

An AI-visibility number is only meaningful if the next number is comparable, and comparability requires locking: the same queries, the same prompt, versioned, run on a schedule. We built our own tracker this way — eight locked buyer questions, a fixed prompt version, every answer stored verbatim, every score recomputable from the stored text — and its first two runs scored Magrios 0/100. We recorded both and publish the method, because a benchmark you would only publish when it flatters you is not a benchmark.

Why must the benchmark be locked?

Because a query set you can edit after seeing results is a machine for manufacturing improvement. Drop the question you lost, add one you win, and the trend line obeys. Our rule is mechanical: the tracked queries are fixed, additions go at the end, and removing one is a founder-level decision precisely because it breaks run-over-run comparability. The locked benchmark methodology is the general argument; this article is the working instance.

The prompt is part of the lock. Ours is versioned, and it deliberately adds no framing that would coax a brand mention — the query is passed as a buyer would ask it, and what the model volunteers is the measurement. Reword the prompt to hint at your brand and you are measuring your hint.

The queries themselves must be unbranded category questions — the questions buyers ask before they know you exist. Branded queries are the wrong benchmark: asking a model about your own name measures spelling, not standing.

What exactly does the score measure?

Each stored answer is scored out of 100 from three observable components:

| Component | Weight | What it detects |

|---|---|---|

| Presence | 50 | The answer mentions the brand at all — absence from the answer is absence from the shortlist |

| Citation | 30 | The answer contains the brand's domain — the model can send the buyer to you, not just name you |

| Position | 20 | Where the brand lands among all tracked entities by first appearance: 1st = 20, 2nd = 12, 3rd = 6, later = 2 |

A run's aggregate is the arithmetic mean of scored rows. Failed calls become error rows and are excluded; a run where every call failed has no score at all rather than a misleading zero. And the honest scope statement, printed in the interface itself: a plain chat model answers from training data, so this measures model-knowledge presence — not live answer-engine retrieval. The absolute number is soft; the movement between locked runs is the signal.

Why store answers verbatim?

Because a score you cannot recompute is a score you have to take on faith. Every model answer is stored as returned, and every score is computed from that stored text by pure functions — anyone can re-derive the number, and no one can fabricate it. A provider failure stores the error message, never a guessed row. Mock rows for previewing the interface are stamped as mock in the data and are never aggregated with real measurement. The stored answers are the evidence trail; the score is merely arithmetic on top of it.

What do you do with the number?

Not worship it. A single-digit score on a young company is expected; ours is zero. The operational output is the blind-spot queue derived from the same rows: queries where no provider mentioned the brand while competitors were named (critical), where nobody was named at all (an open answer, high), or where the brand appeared but never first (medium). Those severities are defined by counting, not judgment, and error-only queries are excluded — unmeasured is not a finding. Each blind spot names the queries and the competitors who own them, which is a work order, not a mood.

What are the limits of one run?

One call per provider per query is a sample, and models are stochastic — sampling error is real and a single run's absolute score should be held loosely. The lock is what redeems this: identical questions, identical prompt, on a defined measurement window, so noise averages out across runs while real change accumulates. Our 0/100 baseline may bounce; what we will trust is the locked trend, recorded either way.

Frequently asked questions

What is a locked benchmark in AI visibility measurement?

A locked benchmark fixes the query set and the prompt between measurement runs, so any movement in the score reflects real change rather than a changed test. Additions go at the end of the set; removals are a deliberate governance decision. Without the lock, a team can manufacture improvement by quietly swapping questions it loses for questions it wins.

Why is an AI visibility score out of one run unreliable on its own?

Models are stochastic and one call per provider per query is a small sample, so a single absolute score carries sampling error. The locked design compensates: identical questions and an identical prompt on a schedule make the trend trustworthy even where one number is soft. Magrios's own interface states this scope plainly rather than overselling the score.

Should a failed measurement run count as a zero score?

No. Zero is a finding — the systems were asked and the brand was absent. A failed run measured nothing, so it gets no score at all. Magrios's tracker stores each provider failure as an error row, excludes error rows from the aggregate, and reports an all-error run as null, because a fake zero would corrupt the trend the lock exists to protect.

Further reading — chosen for this article
Entities in this research
Magrioslocked benchmarkAI visibilityLLM rank trackingblind spotsmeasurement integrity
Related knowledge

Measurement questions: what buyers want proven before they pay · linked

Benchmark questions: how buyers calibrate what good looks like · linked

What is a prompt persona? A practical definition · shared entities

Why an all-error run scores null: honest-null benchmark design · shared entities

Recently updated

Magrios vs Athena · 2026-07-22

What is AI share of voice? A practical definition · 2026-07-22

What is Citation surface? A practical definition · 2026-07-22

Magrios vs Writesonic · 2026-07-22

Where does your brand stand?
Check your AI visibility free — real evidence, not a score.
Check my visibility or run the full analysis →