Magrios / Knowledge / AI Visibility / A working confidence taxonomy: measured, derived

A working confidence taxonomy: measured, derived, hypothesis

Guide · AI Visibility · 3 min read · last verified 2026-07-22

Reviewed before publication Editorial board Independent commercial review
In shortThree levels — measured, derived, hypothesis — each defined by the basis of the claim and what would invalidate it, with 'assumed' banned as the leak where fabrication starts. Worked examples from Magrios's own recorded measurements.

A confidence label is only useful if it names the basis of the claim. We run every claim in our system through three levels: measured — directly observed in a run anyone can re-execute; derived — computed from measurements by a stated rule; hypothesis — a proposed mechanism that fits observable behavior but has not been instrumented. A claim that cannot state its basis does not get a softer label. It gets removed.

What does each level mean in practice?

| Level | Basis | Working example from our own system | What would invalidate it |

|---|---|---|---|

| Measured | Direct observation, stored, recomputable | Our AI-visibility baseline scored 0/100 — model answers stored verbatim, score recomputable by pure functions from the stored text | A re-run of the same locked benchmark producing different stored answers |

| Derived | A stated rule applied to measurements | A blind spot marked "critical" — defined mechanically as: at least one provider answered, zero mentioned us, competitors were named instead | Either the underlying rows changing or the rule being shown to misclassify |

| Hypothesis | A mechanism consistent with observable behavior, not yet instrumented | "Answer engines require third-party corroboration before asserting an entity exists" — fits our measured denial, mechanism unverified | An instrumented test, or the observable pattern failing to hold |

The levels are about basis, not about how strongly we feel. A measured claim can be trivial; a hypothesis can be the most important sentence in a report. What the label buys the reader is knowing exactly what kind of pushback would change our mind — and confidence with a basis is the fuller treatment of the first two levels.

Why is "assumed" banned as a confidence level?

Because an assumption is a claim that skipped the taxonomy. "We assume enterprise buyers check citations" sounds like modest hedging, but it launders an unexamined belief into the reasoning chain with no stated basis and no falsifier. Under the taxonomy it must become one of two honest things: a hypothesis, stated as one, with the observable behavior that motivates it — or a deletion. There is no third state where a claim gets to influence a conclusion without declaring what it stands on.

This is not pedantry; it is where fabrication enters serious documents. Almost no one invents a statistic outright. They assume, then the assumption is repeated without its qualifier, then it is cited. The taxonomy blocks step one.

How do claims move between levels?

Downward is automatic; upward costs instrumentation.

What does the taxonomy look like in a shipped report?

Three behaviors a reader can check rather than trust. Claims carry their confidence with its basis, so "derived" always answers derived from what, by what rule. When something is not publicly knowable, the report prints "no public evidence found" instead of a guess dressed as a finding — refusal is an output, not a failure. And every measured claim links back to the stored evidence behind it, because a label is only as honest as the evidence trail under it.

The test we apply, and the one worth applying to any vendor's report: pick any confident sentence and ask which of the three levels it claims, what its stated basis is, and what would prove it wrong. A system that cannot answer in one sentence is not being careful about confidence — it is being careful about appearances.

Frequently asked questions

What is the difference between a measured and a derived claim?

A measured claim is a direct observation with stored, recomputable evidence — a benchmark run whose answers are kept verbatim, for example. A derived claim is computed from measurements by a stated rule, like a blind-spot severity defined by counting mentions. Derivations inherit the fragility of both their inputs and their rule, which is why the rule must be stated.

Why is 'assumed' not a valid confidence level?

Because an assumption is a claim that declares no basis and no falsifier, yet still shapes conclusions. Honest options are two: restate it as a hypothesis grounded in some observable behavior, or delete it. Most fabrication in serious documents starts as an assumption that later loses its qualifier — banning the label blocks that first step.

Can a hypothesis be worth publishing?

Yes — often it is the most important sentence in the analysis, provided it is labeled as a hypothesis and tied to the observable behavior that motivates it. What is not publishable is a hypothesis wearing the costume of a fact. The label tells readers exactly what kind of evidence would change the conclusion, which is what makes the claim usable.

Further reading — chosen for this article
Entities in this research
Magriosconfidence taxonomymeasured claimsderived claimshypothesisevidence basisdecision traceability
Related knowledge

Evidence-chain architecture: from source to recommendation · shared entities

Risk questions: what buyers fear and how evidence answers it · linked

How AI search engines choose their sources — and what it means for your brand · shared entities

AI visibility for B2B SaaS: what buyers research before choosing software · shared entities

AI visibility for ecommerce brands: how buyers research before they buy · shared entities

Recently updated

Magrios vs Athena · 2026-07-22

What is AI share of voice? A practical definition · 2026-07-22

What is Citation surface? A practical definition · 2026-07-22

Magrios vs Writesonic · 2026-07-22

Where does your brand stand?
Check your AI visibility free — real evidence, not a score.
Check my visibility or run the full analysis →