How to set an AI visibility baseline
Guide · Continuous Intelligence · 4 min read · last verified 2026-07-25
A baseline is the difference between measuring AI visibility and guessing at it. Without one, every scan is a snapshot with nothing to compare it against — you see that a competitor appears in ChatGPT's answer today, but you cannot say whether that is new, whether you slipped, or whether the model simply phrased things differently this week. A baseline turns a pile of observations into a trend line. It is the first thing to build and, done right, the last thing you should ever quietly change. This is the practical procedure for setting one.
Start with the questions, not the score
The instinct is to open a tool, run a scan, and treat whatever number comes back as the baseline. That skips the only step that matters. A baseline is not a score — it is a fixed set of the real questions buyers ask before they choose in your market, plus the answers AI assistants currently give to them. Get the questions wrong and every reading afterward measures the wrong thing precisely.
Build the question set from evidence, not from a keyword tool. Pull the phrasings buyers actually use from sales-call notes, won-and-lost deal reviews, support tickets, and the wording of the prompts your own team types into ChatGPT when they are pretending to be a buyer. Cover the full arc: category-definition questions ("what is X software"), comparison questions ("X vs Y"), selection questions ("best X for enterprises"), and objection questions ("is X worth it"). Twenty to forty well-chosen questions beat two hundred generic ones. In Magrios these become the buyer-question map, and the map — not the headline number — is the asset you are actually building.
Fix the field: competitors, assistants, and regions
A baseline compares like with like, so decide upfront what "like" includes. Name the competitors you will track by exact brand string, including the ones you lose to and the substitutes buyers mention instead of you, not just the rivals your marketing deck lists. If you leave out the challenger who keeps surfacing in answers, your baseline will look calm while the market moves underneath it.
Decide which assistants count. Visibility fragments across ChatGPT, Perplexity, Claude, Gemini, Copilot, DeepSeek, Grok, and Google AI Overviews, and a brand strong in one can be absent from another. Pick the surfaces your buyers actually use and commit to them; adding a new assistant later changes the composition of the score, so treat that as a versioned change, not a silent one. Do the same for language and region if you sell across markets — an answer in German pulls from different sources than the same question in English.
Capture the first scan and record the evidence
Now run the questions against the chosen assistants and record three things for every result: whether your brand appears, in what position or framing, and — this is the part people skip — which sources the assistant cited to get there. A baseline that stores only "present or absent" tells you where you stand but nothing about why. The cited-sources view is what later lets you act, because it shows whether the answer leaned on a review platform, a Reddit thread, a competitor's comparison page, or your own documentation.
Store the raw evidence, not just the summary. Keep the URLs, the date, and the assistant version where available. The point of a baseline is that a skeptic — a CMO, a board member, a competitor's champion — can retrace exactly how the reading was produced. Evidence you cannot reproduce is an opinion with a number attached.
Lock it — and decide what "locked" forbids
The moment the first full scan is captured, the baseline is set: freeze the question set, the competitor list, and the assistant roster. Locking is what makes the next reading honest. If the questions drift between scans, you can never tell whether a visibility gain came from better content or from a friendlier question sneaking into the set.
Be explicit about what locking forbids and what it permits. It forbids editing, dropping, or rephrasing baseline questions to flatter a result. It permits adding new questions — but only as a separate, parallel set that runs alongside the original and never overwrites it. Write this rule down where the whole team can see it, because the pressure to "just tweak one question" always arrives the first time a scan looks bad.
Set the cadence and the first re-scan
A baseline is a starting line, not a monument. Decide how often you will re-run it against the locked questions — monthly suits most B2B categories, faster around a launch or a competitor's funding event — and put the re-scan on the calendar before you need it. The first re-scan is where the baseline earns its keep: it produces your first real delta, the first evidence-backed statement that you gained, held, or lost ground on questions that matter.
From there the work is a loop, not a report. Read the deltas against the locked line, route the biggest movements into concrete actions on the sources AI actually cites, and let the next re-scan judge whether the action worked. Magrios exists to run that loop continuously against a benchmark that does not move under you — because a baseline you keep adjusting is just a mirror, and the whole reason to set one is to stop flattering yourself and start measuring.