How do I measure my brand’s visibility in AI search answers
Guide · AI Visibility · 5 min read · last verified 2026-07-21
To measure your brand's visibility in AI search answers, do three things: fix a benchmark of real buyer questions; measure whether your brand is present in the answers to those questions and in the sources those answers draw on, stating clearly which of the two you measured; and re-measure the same benchmark on a regular cadence. The metrics that matter are presence on the locked question set, the share of questions you win versus who appears instead, and citation of your own domain in the source layer. A single run — or any run against a changed question set — is a snapshot, not a measurement.
That is the whole method. Each step below has a tempting shortcut that quietly destroys the measurement, so the rest of this piece is about doing them honestly.
Step one: build a benchmark that means something
The benchmark is a fixed list of questions your real buyers ask, in the phrasing they use — drawn from sales calls, support tickets, and search query data. Not the questions you wish they asked, and not prompts engineered so your brand appears. If buyers ask "how do I keep customer data out of vendor tools," that belongs on the list even if answers to it never mention your category by name; that is where the buying decision forms. A few dozen questions is a workable starting size.
Two properties make the set trustworthy. It must span intent stages: definitional, comparison, and late buying questions behave differently in AI answers, and a set skewed to one stage will mislead you. And it must lock. Once measurement starts, you do not add, remove, or reword questions; you version the set, and a new version starts a new trend line. A locked set is what makes movement meaningful — an edited set makes every trend unfalsifiable, because improvement might be the questions changing rather than your visibility. If answer engine optimization (AEO) is the discipline of earning presence in AI answers, the locked benchmark is how you know whether it is working.
Step two: record the right things per question
For every question, record four things — and nothing you cannot record again the same way next time.
Present or absent. Does your brand appear? Be explicit about the layer: in the generated answer, in the cited sources, or both. These are different facts — source-level presence is more stable and more diagnosable; answer-level presence is what the buyer sees. Recording them separately lets you later explain a change instead of merely observing it.
Who appears instead. Absence alone tells you almost nothing. Absence while competitors are present means the engines consider the question answerable and have chosen sources — just not yours. Absence while no vendor appears means the question is being answered generically, a different problem with a different fix.
The source pages. Note the pages answers draw on: which domains, which page types — comparison pages, documentation, community threads, glossaries. This is the layer you can act on. You cannot edit a model's answer, but you can influence, correct, or compete with the pages it reads.
Confidence. AI answers vary between runs of the same question. Record how you sampled — how many runs, which engines, what date — and resist reading meaning into differences smaller than the variation you saw between runs. A measurement without a stated confidence is an anecdote wearing a number.
Step three: re-measure the same set, on a cadence
Pick a cadence you can sustain — monthly works for most teams — and re-run the identical benchmark the identical way: same questions, same sampling, same recording format. The value compounds; the third measurement tells you more than the first two combined, because only then can you separate movement from noise.
Read trends honestly. Declines are data — often the most valuable data, because they surface a competitor's content displacing yours or a source page decaying while nobody watched. And a changed question set is not a trend: alter the benchmark and you have started a new measurement, so comparing across the change is self-deception. This is also the core reason one-off AI visibility audits mislead: a single sample of a variable system, however carefully taken, cannot distinguish position from luck.
The metrics that matter — and the vanity ones
Four metrics carry decision weight: presence rate across the locked set; questions won, where you appear rather than a competitor; citation of your own domain in the source layer; and movement on the same locked set over time. Each is falsifiable, repeatable, and tied to an action you can take.
Vanity metrics are the ones that can rise while nothing improves: total mentions across an unstated question list, single-run visibility scores with no sampling method, follower-style counts of how often a model can be coaxed into saying your name. The test is simple: if a skeptical colleague could not re-measure it the same way, it is decoration.
When a spreadsheet genuinely suffices
A spreadsheet honestly suffices when the question set is small, you track one market in one language, the cadence is monthly, and one person owns the process. Rows are questions, columns are dates, cells record presence, competitors, and sources. Run it that way for a quarter before deciding you need more; the discipline matters more than the tooling.
Tooling earns its place when scale breaks the spreadsheet: source-layer measurement across many pages, variance handling across engines and runs, several markets, or results other people must audit. Then the honest question is not which dashboard looks best but whether the system preserves the method — locked benchmarks, stated scope, traceable evidence. That is the difference between a reporting tool and an Intelligence Operating System: one displays numbers, the other keeps the chain from question to evidence to conclusion intact.
Where Magrios fits
Magrios runs exactly this method as a continuous system: benchmarks built from real buyer questions, presence measured at the source layer as well as the answer layer, re-measurement on cadence against the locked set, and every conclusion traceable to its evidence. For a fast first read before committing to any process, the free AI visibility checker shows how your brand currently appears for a handful of questions. It is a snapshot — with all the limits described above — but an honest way to learn whether you have a measurement problem worth taking seriously.