Which tools track how ChatGPT and Perplexity describe my company
Guide · AI Visibility · 4 min read · last verified 2026-07-21
Two kinds of tools can track how ChatGPT and Perplexity describe your company, and they work in fundamentally different ways. The first is prompt-sampling: asking the assistants a fixed panel of questions on a schedule and recording what comes back. The second is source-layer measurement: measuring whether your company is present, and correctly described, in the public pages assistants draw on when they compose answers. These are complementary instruments that answer different questions, and the single most important property of any tracker is that it tells you plainly which of the two it actually does.
Magrios does not publish ranked lists of trackers it has not evidence-audited, so you will not find vendor names below. Rankings with undisclosed methodology are precisely the content disease this category exists to fight. What follows instead is what each mechanism can honestly claim, the criteria that separate a measurement instrument from a dashboard, and a check you can run this month without buying anything.
What prompt-sampling can and cannot tell you
Prompt-sampling tools query assistants directly, so what they capture is close to what a real user might see. That is their strength, and it comes with structural limits that follow from how assistants work rather than from any vendor's execution.
Assistant answers vary run to run. The same question, asked twice in a row, can produce different phrasings, different competitor mentions, and different cited sources. Answers also vary by user: context, conversation history, location, and model version all shape output, and no external tool can observe the private answers your actual buyers receive. And there is no stable benchmark underneath: when a sampled score moves, the cause could be a model update, a sampling change, or a genuine change in how you are represented — the number alone cannot say which.
None of this makes sampling useless. It can tell you whether you appear at all across a designed panel of questions, roughly how you are framed when you do, and whether a damaging factual error is circulating. But turning samples into trends requires careful design — fixed question panels, repeated runs, logged model versions — and even then the result is an estimate of a moving distribution, not a measurement of a stable quantity. That distinction is the core of why AI chat alone cannot carry strategic decisions.
What source-layer measurement can and cannot tell you
Assistants do not compose answers from nothing. When asked about companies and categories, they draw heavily on retrievable public sources: documentation, comparison pages, reviews, directories, and articles. Source-layer measurement works one level down — it measures whether you are present and accurately described in that evidence layer.
This approach has properties sampling cannot offer. Pages do not vary run to run, so the measurement is stable. Every claim can link to the URL it came from, so findings are checkable. And because the same corpus can be re-measured on a schedule, movement over time is a real trend rather than sampling noise.
The honest limitation is equally structural: source-layer measurement observes the inputs to an answer, not the answer itself. It cannot tell you what one assistant said to one user on one afternoon. A trustworthy system states which layer it measures instead of blurring the two. For how the layers connect — and why presence in the source layer is the part you can actually act on — see AI visibility: the complete guide.
Evaluation criteria for any tracker
Whichever mechanism a tool uses, five criteria separate measurement instruments from dashboards:
- A stable benchmark. Can the same thing be measured the same way next month, so that change means something?
- Per-claim sources. Does every finding link to the evidence behind it, or are you asked to trust a score?
- An honest scope statement. Does the tool say plainly whether it measures sampled answers or the source layer — and what it cannot see?
- A re-measurement cadence. A one-off audit describes a moment; representation problems are ongoing.
- An action loop. When a reading is bad, does the tool point to the specific page or absence causing it, or only report the symptom?
Note also what this category is not: survey-based brand tracking measures human perception, which is a different quantity requiring different instruments. The distinction matters when you are comparing scopes and quotes, and it is covered in brand tracking vs AI visibility tracking.
A DIY monthly check that needs no tools
You can build a usable baseline yourself in about an hour a month. Write ten questions a real buyer would ask in your category — about the problems you solve, not your brand name. Ask each one in a fresh session of each assistant you care about, and record three things: whether you were mentioned, how you were framed, and which sources the answer cited. Then open those cited sources and read what they say about you; this step usually explains the answers. Repeat with the same questions next month, logging the date and model version. Treat single-run changes as noise, and treat patterns that persist across months and across cited sources as signal worth acting on.
Where Magrios fits
Magrios is a source-layer measurement system. It measures whether your company is present, and accurately described, in the public evidence assistants draw on, links every claim to its source, and re-measures on a cadence so that movement reads as trend rather than noise. The explicit limitation: Magrios does not measure private assistant answers — it cannot tell you what ChatGPT told a specific user yesterday. If that is the question you need answered, run the sampling check above alongside it. What Magrios tells you is whether the public record those answers are built from is one you would want repeated.