The locked benchmark: why honest measurement locks its questions
Guide · AI Visibility · 5 min read · last verified 2026-07-22
A locked benchmark is a fixed set of real buyer questions whose membership cannot silently change between measurements. Because every scan asks exactly the same questions in the same form, any movement in the results is market movement — a competitor gaining ground, an answer engine changing its sources — rather than measurement movement. In market intelligence, and in AI visibility tracking especially, locking the question set is what separates a trend you can act on from a sequence of unrelated snapshots dressed up as one.
Why an unlocked benchmark cannot produce a trend
A trend is a claim that the same thing was measured twice. If the thing being measured changes between readings, the claim dissolves — quietly, because the chart still renders and the line still moves. When a report says your visibility in AI answers improved between two scans, the first question worth asking is whether both scans asked identical questions under identical rules. If they did not, the improvement may be an artifact of the question set, and no amount of care downstream can repair it. The mechanics of running such a measurement — which questions, which assistants, what counts as presence — are covered in How do I measure my brand's visibility in AI search answers; locking is the discipline that makes the second run mean something relative to the first.
The three ways unlocked benchmarks lie
Silent question swaps. Questions get replaced between scans, often for defensible reasons. But if the replacement is silent, every comparison spanning it is corrupted invisibly. A brand can appear to improve purely because questions where it was absent were swapped for questions where it shows up.
Methodology drift. The questions stay but the asking changes: different prompt phrasing, different assistant versions treated as equivalent, a looser rule for what counts as a mention. Drift is more dangerous than swaps because nothing in the question list betrays it; the benchmark looks identical while measuring something else.
Survivorship pruning. The most tempting failure: quietly dropping questions where the client looks bad. Each removal is defensible alone, and the cumulative effect is a benchmark structurally incapable of reporting bad news. A pruned benchmark does not merely err; it errs in one direction, always flattering.
The locking rules
Three rules are enough to close all three failure modes:
- Additions only at the end, and only versioned. The benchmark may grow, but new questions are appended in a new, named version. No question changes position, wording, or identity inside an existing version.
- Removals never. A question that has genuinely stopped mattering can be flagged and excluded from headline reporting, but it stays in the set and keeps being measured. Deleting it would delete its history, and its history is the record.
- Methodology changes restart the baseline. When the measurement procedure itself changes — new assistants in scope, new mention criteria — the honest response is a fresh baseline, clearly marked, not a stitched chart pretending continuity.
These rules are what make continuous measurement more than a subscription to repeated one-off audits. The broader argument for continuity over snapshots is made in Static market reports vs continuous intelligence; locking is the mechanism that keeps the continuous version honest.
Deep dive: comparability adjudication
The mature form of locking is not a policy document but an adjudicator: a rule system that decides, for any two measurements, whether comparing them is permitted — and refuses when it is not.
Two scans may be honestly compared when all of the following hold: they ran against the same benchmark version; both runs completed; and both used the same methodology version, including the same set of assistants in scope.
A comparison must be refused, with the reason stated, when any of these fail:
- Different benchmark versions. Even one appended question shifts every aggregate. The defensible fallback is a labelled core comparison: score both scans on the older version's questions only, and say so explicitly. Worked reasoning: if version two appends questions to version one, comparing a v1 scan with a v2 scan on the full v2 set is refused outright; on the v1 core it is allowed, but the label — core comparison, v1 questions only — must travel with the number wherever it goes.
- Incomplete runs. If a scan failed partway, comparing it to a complete scan converts an infrastructure failure into a fake market signal: absent answers read as lost visibility.
- Different methodologies. Cross-methodology deltas are not adjustable after the fact; they are simply not deltas. The adjudicator points at the re-baseline instead.
Refusal is the trust-building move, not the embarrassing one. Every refused comparison is an assertion the system declined to manufacture, and a tool that sometimes says these two numbers cannot be compared, and here is why, is demonstrating the property buyers actually need from measurement. This is the evidence-first standard applied to the measuring instrument itself — a claim either carries its verifiable grounds or it is not made, as argued in Evidence-first AI: what it means and how to verify a vendor's claim to it.
Decline reporting: the credibility test
A benchmark that only ever reports gains is an advertisement wearing a lab coat. The strongest single signal that a measurement system is honest is that it can say you lost ground — computed under the same locked set and the same rules as the gains. Locking is what makes a decline useful as well as believable: because the questions are fixed, a drop resolves to specific questions where a specific answer changed, which is an investigable event rather than a mood. A system that has never once reported a decline has told you nothing about the market and everything about itself.
When locking hurts
Equal candour requires admitting that locking has a real cost. Markets pivot. Buyers stop asking last year's questions, and a faithfully locked set can end up measuring, with perfect comparability, a conversation that no longer exists. Locking-zealotry produces benchmarks that are precisely consistent and increasingly irrelevant.
The resolution is deliberate re-baselining, not quiet editing. Build the new question set for the market as it now is; run it alongside the old one for an overlap period so the discontinuity is visible; then archive the old benchmark with its history intact and mark the break on every chart that spans it. What honesty forbids is not change — it is silent change. An edited benchmark and a re-baselined one may end up containing the same questions; only one of them can be audited.