Blog · 2026-07-20 · 6 min read
The grounding rule: most AI-visibility numbers measure the wrong thing
Ask a model a question with web search off and you learn what it memorized in training. Ask with search on and you learn what a buyer actually sees. Confusing the two is the original sin of AEO measurement.
Here's an uncomfortable question to put to any AI-visibility number: how was the answer generated? Not which model, not how many prompts — was the engine's live web search on or off when it answered?
It sounds like a plumbing detail. It's the whole measurement. A large language model answering from its training data is reciting a snapshot of the web as it existed months or years ago, compressed and blended. A model answering with live search enabled is doing what it does for a real buyer: retrieving current pages, reading them, and composing an answer with citations. These are two different systems that happen to share a chat box.
Ungrounded numbers flatter incumbents
Training data over-represents the past. A brand with fifteen years of accumulated web presence looks strong in an ungrounded answer even if its live visibility is collapsing; a two-year-old challenger doing everything right can look invisible. If your measurement is ungrounded, you're not tracking the market — you're tracking history.
Grounded answers move with the web. Publish the page, get into the roundup, fix the schema — and a grounded answer can reflect it in the next run. That's what makes grounded measurement actionable: it responds to the work you do. An ungrounded score barely can, because the training data won't update for months, if ever.
This is why every SolvedAgain query runs with live web search on, and why every stored answer carries a grounded flag. When an engine can't ground, the report says so — the flag exists precisely so a number is never presented as something it isn't.
The second rule: sample, because answers vary
Grounding is necessary and not sufficient. AI answers are probabilistic — the same grounded question, asked twice, can name different brands in a different order. A single sample is an anecdote wearing a percentage.
The honest fix is boring and mechanical: ask every question multiple times per engine — we run three — store every answer, and report the disagreement instead of averaging it away. We surface that as a volatility metric. High volatility on a question isn't noise to hide; it's information. It tells you the engine hasn't settled on an answer, which usually means the category is winnable.
Sampling also keeps you from fooling yourself in the good direction. One run where the engine names you first feels like victory; three runs where you appear once feels like what it is — a coin toss you're starting to influence.
Questions to ask any vendor — including us
Was web search on? Per answer, provably? Can I read the raw answers my score came from? How many samples per question? What happens when the samples disagree? If any answer is a shrug, the number on the dashboard is a vibe with a decimal point.
Our answers, for the record: always on and flagged per answer; yes, every number traces to a stored answer you can open; three per engine; disagreement is reported as volatility, never averaged away. Measurement you'd bet budget on has to survive its own methodology page.
— The SolvedAgain team