How to choose a marketing agency in the AI era
Guide · Enterprise · 4 min read · last verified 2026-07-27
Choosing a marketing agency is a hiring decision made with less information than almost any hire: you are buying judgment you have not watched being exercised, on the strength of a pitch engineered to prevent exactly that inspection. The AI era has not changed the shape of that problem, but it has changed which questions separate firms, because AI has changed what agencies do all day — and the gap between shops that use it as leverage and shops that use it as a margin trick has become, in our experience, one of the clearest quality divides in the market.
The questions that changed
Three questions now do most of the separating work.
First: is AI leverage or dilution inside their shop? Every agency will claim thoughtful adoption, so ignore the claim and ask for the mechanics. The distinguishing answer describes where AI drafts and where senior people decide, what is never allowed to ship unedited, and how the firm keeps assisted output from converging on the generic. Vague talk of efficiency, with no named checkpoint where judgment intervenes, tends to precede volume-shaped work arriving in your brand's name.
Second: do they measure position or activity? Deliverables shipped, impressions served, and posts published are activity. Position is what your buyers actually encounter when they look — including, now, what AI assistants say when asked the questions your buyers ask, which in current examples varies by engine and shifts over time. An agency that proposes to be judged on movement against a locked set of buyer questions is volunteering for accountability; one that reports effort is engineering an escape from it. Whether they run their own tooling or a platform like Magrios, insist that the question set and its history live in a workspace you own, so the baseline outlives the relationship.
Third: will they show the evidence behind recommendations? Strategy decks are cheap to make impressive. The valuable habit is a trail — this recommendation, because of this observation, sourced here, checkable by you. Firms with that habit tend to answer evidence questions quickly and specifically; firms without it fill the silence with anecdote.
Evidence tests you can run in the room
Ask for one past recommendation and the trail behind it, anonymized however they need. You are not judging whether the call was right; you are judging whether a chain from evidence to advice exists at all.
Ask what they have advised a client not to do. Agencies bill for activity, so a firm that can name revenue it argued against — a channel abandoned, a campaign killed, a trend declined — is showing you its judgment operating against its own short-term interest.
Ask for a call they got wrong and what changed afterward. Everyone has one; the tell is whether the story ends with a process change or with an excuse.
Then run one small live test: give them a real buyer question from your category and ask how they would establish where you stand today. Listen for method — sources, cadence, evidence, cost of being wrong — rather than a recitation of deliverables.
Reference checks that actually work
References are curated, so the questions must do the work the curation undoes. Ask what the agency stopped doing once it clearly was not working, and how long that took — persistence with a failing tactic is the expensive failure mode of retained relationships. Ask how the agency behaved when it missed: did the client hear it from the agency first, or discover it themselves? Ask who did the work after the pitch team disappeared, because the distance between the sellers and the doers tends to predict your experience. And ask what the client would not hire this agency for again; honest references usually have an answer, and it maps the edge of the firm's real competence better than any case study.
Red flags that end the conversation
Guaranteed rankings, guaranteed placement in AI answers, guaranteed citation: walk. Engine behaviour is observed, not controlled, and it changes; a firm guaranteeing outcomes it cannot control is telling you what its other promises are worth. Volume-first AI proposals — content at scale as the opening move — signal dilution sold as strategy. A methodology too proprietary to show you its evidence is a wall, not a moat. Reporting that never arrives at a decision recommendation is activity theater. And any arrangement where your measurement history, accounts, or question set would leave with the agency is a dependency sold as a service.
Running the selection end to end
Write the brief before the shortlist: the decisions you need moved, the questions you want owned, the evidence standard you expect. Shortlist firms whose published thinking shows judgment rather than volume. Run the in-room evidence tests, then a small paid pilot scoped to a handful of named decisions — paying for pilots keeps both sides honest and keeps your claim to the work clean. Establish your own baseline before the pilot starts, so every firm is measured against the same yardstick rather than against its own reporting. Then decide the way you would for a senior hire: slowly on judgment, quickly on integrity. An agency that resists being measured during courtship will not become more measurable after the contract is signed.