aeo · measurement

The fixed-panel method: measuring AI visibility month over month

Three identical instrument panels in a dark rack, each round screen showing the same waveform, with the third screen glowing brightly

Here's a scene that plays out in a thousand businesses a week: an owner asks ChatGPT "who's the best [their trade] in [their city]," sees a competitor's name, and panics. Or sees their own name and celebrates. Both reactions are built on nothing — ask again tomorrow, in a different session, and the answer may change. A single query is an anecdote. AI visibility only becomes manageable when you measure it like a metric — and metrics require consistency.

This is the method I use to do that. I run it monthly on client accounts and on my own e-commerce brand, where I've watched the numbers move as entity work landed. It's deliberately simple, because a measurement system only survives if it's easy to repeat.

The problem with checking "sometimes"

Large language models are probabilistic. The same question can produce different answers across sessions, and the assistants themselves change constantly — new model versions, new retrieval behavior, new interfaces. If you ask different questions, at different times, in different ways, you have no baseline. You can't tell whether last month's schema work moved anything or whether you just caught the model in a different mood. The fix is the same one survey researchers landed on decades ago: hold the instrument constant and let the world vary.

The method, in five rules

  • Lock the panel. Choose a small, fixed set of queries — three to ten — that represent how real customers ask. Mix intent types: a direct hire question ("who should I call for … in …"), a comparison ("best … near …"), and a category or product question. Write them down. The exact wording never changes again.
  • Fix the assistant set. Pick the surfaces that matter for your market — for most local businesses that's ChatGPT, Gemini, and Google's AI Overviews — and use the same set every month. Adding a new assistant later is fine; it starts its own trend line.
  • Run each query more than once. Three runs per query, in fresh sessions, smooths the randomness. What you record is the pattern across runs, not the outcome of one.
  • Record structured fields, not screenshots. For every run: Was the brand named? Was its site cited as a source? Which competitors were named? What sources did the assistant cite? A spreadsheet with those four columns beats a folder of screenshots you'll never open.
  • Same time, every month. Cadence is the whole point. One measurement is trivia; twelve are a story.

What the scoreboard looks like

Roll the runs up into three numbers per assistant: mention rate (how often you're named), citation rate (how often your site is the quoted source), and competitor share (who else keeps showing up). An illustrative month for a fictional plumber:

AssistantMention rateCitation rateTop competitor named
ChatGPT6 of 9 runs3 of 9 runsWestside Plumbing
Gemini4 of 9 runs2 of 9 runsWestside Plumbing
AI Overviews7 of 9 runs5 of 9 runs

Numbers like these do two jobs. They tell you where you stand — and, tracked over months, they tell you whether the work is working. When a citation rate climbs after you've rewritten a page to be quotable, that's not a vibe. That's attribution.

The competitor column earns its keep fast. When one rival keeps appearing in answers you're missing from, their footprint — the pages they've published, the sources citing them, the reviews they're generating — is a map of exactly what the models are rewarding in your market.

Honest limitations

This method measures direction, not absolutes. Variance never fully disappears; personalization means your customers may see slightly different answers than your clean test sessions; and a model update can shift a baseline overnight through no action of yours. None of that is a reason to skip measurement — it's the reason to measure consistently, so you can tell a real trend from a noisy month. Note the date and any known model changes alongside each month's numbers and the trend line stays interpretable.

Why I publish this

Because most of the industry is still treating AI visibility as either unknowable or as a screenshot in a sales deck. It's neither. It's a measurable channel, and the businesses that start measuring now will be the ones who understand their market's answer landscape while competitors are still asking the chatbot once and guessing. If you want the mechanics of what actually moves these numbers, start with how AI decides which local business to recommend.

I run this benchmarking as part of every AI visibility engagement — your panel, your market, your trend line, reported monthly in plain English. Book a call if you want to know what the assistants are currently saying about you.

← Back to the blog