How to run an AI visibility audit in a week
Guide · AI Visibility · 4 min read · last verified 2026-07-25
You do not need a quarter, a consultant, or a new platform contract to find out how visible your brand is in AI answers. You need a focused week and a method that resists wishful thinking. The point of a one-week audit is not to produce a perfect number — it is to replace a vague anxiety ("are we even showing up in ChatGPT?") with a concrete, defensible baseline and a short list of the gaps worth fixing first. Here is how to run that week without fooling yourself.
Day 1 and 2: build the question set, not the keyword list
Everything downstream depends on asking the right questions, so resist the urge to start with keywords. AI visibility is measured against the questions real buyers ask assistants, and those come in recognizable shapes: definition questions, comparison questions, "best tool for X" questions, and questions about specific jobs your product does. Write thirty to fifty of them in plain language, the way a buyer would actually type or speak them, not the way your SEO tool phrases them.
Anchor the set to your category and your buyers' decision moments, and include the uncomfortable ones — the comparisons where a competitor might beat you, the "alternatives to [competitor]" queries, the objection-shaped questions. A question set that only flatters you will produce an audit that only flatters you. Aim for coverage across the funnel rather than depth on your favorite topic.
Day 3: capture what the assistants actually say
Now run the questions through the assistants your buyers use — ChatGPT, Perplexity, Claude, Gemini, Copilot, and Google's AI Overviews at minimum. For each answer, record three things and nothing extra at first: whether your brand appears at all, how it is described when it does, and which sources the assistant cites to support the answer.
Two disciplines keep this honest. Ask each question in a clean session so a previous answer does not contaminate the next, and ask the important ones more than once, because model outputs vary between runs. A single lucky mention is not visibility; a mention that recurs across runs is. Treat one-shot results as anecdotes, not data — sampling variability is the most common way a quick audit lies to the person running it.
Day 4: read the citation surface
The list of sources the assistants cite is the most actionable output of the whole week, and most audits skim past it. Sort every cited source into buckets: your own domain, review platforms, independent media and creators, community sources like Reddit, and competitors' pages. The pattern tells you where the models are actually getting their answers about your category, which is almost never where you assumed.
This is where the gaps become concrete. If comparison questions are answered entirely from review platforms and your presence there is thin, you have found your first priority. If your competitor's blog is cited on the definitional questions and yours is absent, you have found your second. The citation surface converts a fuzzy "we need more content" into a ranked list of specific sources to earn or fix.
Day 5: turn findings into a benchmark and a shortlist
End the week with two artifacts. The first is a baseline you can defend: for this fixed question set, on these dates, across these assistants, here is our presence rate, here is how we are described, and here are the sources doing the talking. Write down the methodology in enough detail that you could reproduce it, because a baseline you cannot reproduce is a number you cannot trust later.
The second is a shortlist of three to five gaps, each tied to an action and a source. Not twenty — three to five, chosen because they sit on high-intent questions where you are absent or mischaracterized. The Princeton GEO study (KDD 2024) offers useful direction on what tends to move answers: content that cites sources saw roughly a 40% visibility lift, statistics about 37%, and quotations about 30%, so gaps you can close with corroborated, evidence-rich content are good candidates to attack first.
From a one-week snapshot to a standing measure
A week gives you a snapshot; the value compounds only when you can compare the next snapshot to this one on identical terms. That is the difference between an audit and a program, and it is what Magrios is built to carry: the same buyer-question set, run on a schedule, scored against a locked benchmark so a real change stands out from ordinary model drift, with each gap routed to an action and re-checked on the following scan. Run the week manually to prove to yourself the picture is worth having. Then hold the method fixed, because a baseline is only as useful as your discipline in never quietly moving the goalposts underneath it.