Magrios / Knowledge / Continuous Intelligence / Branded queries are the wrong benchmark for AI v

Branded queries are the wrong benchmark for AI visibility

Guide · Continuous Intelligence · 5 min read · last verified 2026-07-21

Reviewed before publication Editorial board Independent commercial review
In shortBranded AI prompts test recall for brands the model already knows, not real visibility. Unbranded, category-level prompts are the benchmark that actually discriminates between competitors.

Branded queries — prompts that already name your company — are a weak benchmark because known brands get mentioned almost regardless of AI visibility work, leaving little room for the number to move. Unbranded, category-level prompts test whether you're found before anyone names you.

What a branded query actually tests

A branded query is any prompt that already contains the company's name: "What does Acme do," "Is Acme good for X," "Acme vs Competitor." These prompts test recall, not discovery — they ask the model to retrieve and describe something it already has a name for. That's a fundamentally different task than the one a prospective buyer performs when they don't yet know which vendor to consider.

Recall for an established brand is close to guaranteed once a model has any training signal or retrieval access to that brand at all. The AI doesn't need visibility work to answer "what does Acme do" — it needs the brand to exist and be documented somewhere. That's a low bar most funded, operating companies clear by default, which is exactly why the query makes a poor benchmark: it's not sensitive to whether visibility efforts are working.

This is where a lot of reporting confusion starts. A team invests in visibility work, checks a branded prompt set, and finds the number was already high before any of that work began — because it was never testing the thing the work was meant to change. The metric doesn't fail loudly; it just quietly measures something that was true regardless of effort, which makes genuine progress on the harder problem — being found before anyone names you — invisible in the one number leadership is looking at.

Why the number barely moves

Any benchmark is only useful if the number it produces can move in response to the thing being measured. This connects back to AI share of voice — share of voice compares appearance rate against competitors across a shared set of prompts. If the prompt set is dominated by branded queries, the ceiling sits close to the floor for every established competitor: everyone with a real product scores near the top, because the model already knows all of them by name. That compresses the range the metric can report, which makes it look stable even when a business's actual visibility position is deteriorating in the answers that matter.

The prompts that discriminate are the ones where the model has to make a choice: which vendors to mention when the user hasn't specified one. That's where a real gap between competitors shows up, because it depends on what the model has learned to associate with the category, not just what it has been asked to recall by name.

The benchmark that actually discriminates

Unbranded, category-level prompts — "best tool for X," "how do teams handle Y problem," "alternatives for this kind of problem" without naming a product — put every competitor on equal footing at the starting line. None of them are named in the prompt, so whichever ones the model surfaces reflects something closer to actual visibility: training data prevalence, third-party content, review coverage, and how well a company's positioning has been picked up and repeated across the sources the model draws from.

This connects directly to why absence compounds in AI search: a brand that never appears in unbranded answers isn't losing a single sale, it's failing to enter consideration at all, repeatedly, across every prompt where a buyer hasn't yet decided who to ask about by name. Branded-query performance says nothing about that failure mode, because a branded query, by definition, only gets asked by someone who already knows the brand exists.

Worked example: two benchmarks, two different stories

Consider a hypothetical company tracked on two prompt sets. The branded set has 10 prompts, all naming the company directly. Across a sampling pass, the company is mentioned accurately in 9 of those 10 — a 90 percent hit rate that looks excellent on a dashboard.

The unbranded set has 30 prompts describing the problem the company solves, without naming any vendor. Across the same sampling pass, the company appears in only 6 of those 30 responses — a 20 percent hit rate, with two competitors appearing in 22 and 18 of the 30 responses respectively.

Reported alone, the branded number, 90 percent, suggests strong AI visibility. The unbranded number, 20 percent, tells a very different story: most of the buyers asking about the problem category, without already knowing the brand, never see it mentioned at all. Only the second benchmark is sensitive to the actual competitive position.

What to track instead

A branded-query check still has a place — it confirms the model has accurate, current facts about the company when asked directly, which matters for correcting outdated or wrong information. But it shouldn't be the headline metric. The benchmark that reflects real visibility work is built from unbranded, category, and problem-framed prompts, weighted toward the language a buyer would actually use before they know who to ask about by name.

Pairing that unbranded prompt set with a stable group of control questions makes the benchmark even more reliable, because it separates genuine gains in category visibility from noise introduced by model updates or ordinary sampling variation. The goal isn't to stop asking branded questions — it's to stop mistaking a high score on an easy test for evidence that the hard test would pass too.

A useful transition for teams that have only ever tracked branded prompts is to run both sets side by side for a while, rather than swapping one for the other overnight. Watching the two numbers diverge — a flat, high branded score next to a lower, more volatile unbranded one — is often the clearest internal demonstration that the two benchmarks are measuring different things, and that the harder, more honest number is the one worth building a strategy around.

Frequently asked questions

Why do branded AI queries score so much higher than unbranded ones?

Because a branded query only asks the model to recall a company it can already name, which most operating businesses pass by default regardless of visibility work.

Should we stop tracking branded queries altogether?

No — they're still useful for checking that the model has accurate, current facts about your company. They just shouldn't be the main benchmark for visibility.

What makes an unbranded prompt useful for benchmarking?

It doesn't name any vendor, so the model has to choose who to mention, which is where real differences between competitors actually show up.

How many unbranded prompts do we need to get a reliable read?

More than a handful, and ideally paired with control questions and repeat sampling, since any small prompt set carries sampling error on its own.

Further reading — chosen for this article
Entities in this research
branded queryunbranded queryAI visibilityshare of voicelarge language modelcategory promptrecalldiscovery
Related knowledge

What is a prompt persona? A practical definition · shared entities

Adding prompts changes your score without changing your position · linked

Champion vs coach: the distinction that decides whether your deal survives a reorg · shared entities

Recently updated

Magrios vs Athena · 2026-07-21

Magrios vs Writesonic · 2026-07-21

Magrios vs Semrush · 2026-07-21

Magrios vs peec · 2026-07-21

Where does your brand stand?
Check your AI visibility free — real evidence, not a score.
Check my visibility or run the full analysis →