How to run a competitive AI visibility audit
Guide · Continuous Intelligence · 5 min read · last verified 2026-07-25
When a buyer asks an AI assistant "what's the best tool for X," the reply names a short list and quietly drops everyone else. A competitive AI visibility audit exists to answer one blunt question: on the queries that decide your deals, which vendors get named, cited, and described well — and on which of those questions is the winner a rival instead of you?
What a competitive AI visibility audit actually measures
A competitive AI visibility audit measures, question by question, whether an AI assistant surfaces your brand versus a defined set of competitors — as a mention, a citation, or an outright recommendation — across the buyer questions that matter in your market. The output is a presence map, not a single grade. You can act on a map; a grade only tells you how you feel.
The reason to run one is that AI answers are winner-shortlist, not winner-take-all. A rival can own three of your top ten questions while you own none, and a blended "share of voice" number can hide that entirely. The audit's job is to expose the specific questions where you are absent and someone else is present, because those are the gaps that cost pipeline.
Step 1 — Define the question set buyers actually ask
Start with the questions, because everything downstream is scored against them. Pull real buyer language from sales-call notes, your CRM's closed-lost reasons, support tickets, and the "People also ask" and autocomplete surfaces for your category. Group them into intent shapes — comparison ("X vs Y"), selection ("best X for mid-market"), method ("how to do Z"), and risk ("is X compliant with W").
Aim for 30–60 questions in a first pass, weighted toward high-intent comparison and selection queries. Comparison and selection pages tend to earn a disproportionate share of AI citations, so under-representing them will understate where the real battles are fought. Write each question the way a buyer would type or speak it, not the way your marketing team phrases it.
Step 2 — Pick the competitor set (and why fewer is better)
Choose the smallest set that reflects who you actually lose to. Three to six named competitors is usually enough: two or three head-to-head rivals, one category leader, and one insurgent. Adding more inflates the work and dilutes the read. A common mistake is anchoring the whole audit on your loudest competitor, which distorts strategy toward whoever markets hardest rather than whoever wins deals.
Include the "no-brand" case too — questions where the assistant recommends a category or an approach without naming any vendor. Those are green-field questions where the market has no default answer yet, and they are often the cheapest to win.
Step 3 — Capture per-question presence, not a single score
Run each question on each assistant you care about — for example ChatGPT, Perplexity, Google's AI surfaces, and Claude — and record what actually happens, not just yes/no. For every question capture: were you mentioned; were you cited (with a live source link); were you recommended; which competitors appeared; and which sources the answer leaned on.
A compact matrix keeps this legible:
| Buyer question | You | Rival A | Rival B | Top cited source |
|---|---|---|---|---|
| Best X for mid-market | Absent | Recommended | Mentioned | Review site |
| X vs Rival A | Mentioned | Recommended | Absent | Rival A blog |
| How to migrate to X | Cited | Absent | Absent | Your docs |
Because assistants sample and paraphrase, a single run is a snapshot with real run-to-run variance. Run each question a few times, or on a fixed cadence, and treat the pattern as the signal rather than any one answer.
Step 4 — Read the gaps: where rivals are cited and you are not
Now read the map for three failure modes. First, absence: questions where you never appear and a competitor is recommended — the most expensive gap, because you are eliminated before a human evaluates you. Second, misrepresentation: you appear but the description is wrong, thin, or dated. Third, weak corroboration: you are mentioned but the answer's cited sources are all your own pages, which reads as unverified to both the model and the buyer.
For each gap, follow the evidence trail. Look at which sources the assistant actually cited for the questions your rival won — third-party review sites, Reddit threads, documentation, comparison pages. Third-party corroboration tends to be cited more than vendor-owned pages, so a rival winning on independent sources is a different problem than a rival winning on its own blog.
Step 5 — Prioritize the gaps worth closing
Not every gap deserves a fix. Score each on two axes: buyer intent (how close the question sits to a purchase decision) and evidence effort (how hard it is to become the best-corroborated answer). High-intent, low-effort gaps come first. A comparison question you are absent from, where the current top source is a stale forum post, is a far better target than a top-of-funnel definition you already partly own.
Keep the output as a ranked blind-spot queue, not a leaderboard. The number of questions you "win" is vanity; the ordered list of specific absences you are going to close — and then re-check — is the working artifact.
Lock the method so the next audit is comparable
An audit you cannot repeat is an anecdote. Freeze the question set, the competitor set, the assistants, the number of runs, and the scoring rubric, and store that as your benchmark. Change one variable at a time. If you add ten questions next quarter, your aggregate presence will move purely because the denominator changed — not because your position did. A locked methodology is what turns two snapshots into a trend you can trust.
How to operationalize this without doing it by hand
Done manually, a real audit is dozens of questions times several competitors times multiple assistants times multiple runs — and it goes stale the moment answers shift. This is exactly the loop a platform like Magrios runs continuously: a locked question-and-competitor benchmark, per-question presence captured across assistants with the source behind every claim, the biggest absences routed into a prioritized queue, and a re-scan that tells you whether the gap actually closed. The value is not the headline number; it is the repeatable comparison and the ranked list of what to fix next.