SEO & Search
AI Visibility Reports Need Better Evidence, Researchers Warn
Research shows models abandon correct answers 83-93% of the time when tools return conflicting data, raising the bar for AI visibility diagnostics sold to brands.

Marketing teams selling AI visibility diagnostics face an evidence problem, according to a new analysis by Pedro Dias published on Search Engine Journal. Three recent research papers suggest that common explanations for missing brand mentions in AI answers — content gaps, authority problems, recall failures — rarely come with the experimental support to back them up.
The first study, MemToC, tests what happens when a language model's own correct answer conflicts with information returned by an external tool, similar to how retrieval works in RAG systems. Researchers first required models to answer factual questions without tools, then asked again with controlled tool returns. In cases where an instruction-tuned model had answered correctly but the tool returned an incorrect answer, correct-answer retention ranged from just 6.5% to 17.1% across four models, pooled over three instruction wordings.
The model had already given the right answer. That makes "it doesn't know" a poor explanation on its own, Dias argues. But answering a fact correctly once does not show how reliably the model has learned it, or whether the answer will survive conflicting information.
In a separate annotation sample of 120 responses to incorrect tool returns, covering five models and both conflict cases, none of the responses explicitly acknowledged disagreement with the tool. Dias cautions the result is limited to the inspected responses and cannot establish that models never flag conflicts.
The tests were controlled experiments, principally on open-weight models with 7-9 billion parameters. The percentages cannot be applied to ChatGPT search or Google AI Overviews.
Same red cell, different causes
The commercial stakes are straightforward. A visibility report shows a red cell where a brand mention should be. Someone has to explain it at the next client meeting. But a count of appearances tells you only what happened in the answers collected, Dias writes. Calling the red cell an authority problem requires evidence the count does not contain — and it conveniently suggests work to invoice.
A content deficiency and a source conflict can produce the same missing mention. "If the next slide recommends another page, I'd like to know how we settled on a content problem," Dias writes. "Counting the absence again won't answer that."
A second study, "Empty Shelves or Lost Keys?", widens the problem. The authors count a fact as "encoded" if the model reproduces it under either of two strong contextual probes. Their reliable-answering test is stricter: it requires correct answers across all four variants, covering two phrasings and both directions of a factual relationship. That difference in pass criteria helps shape the measured gap.
GPT-5 and Gemini-3 pass the encoding probes for 95-98% of the benchmark's facts, while reliable recall remains weaker. Rare facts and reverse questions are particular problems, though reasoning recovers a substantial share of failures. The study uses Wikipedia-derived facts, so it cannot say how often this happens with commercial brand recommendations.
Dias also draws a distinction practitioners often miss. Ask about a named brand and you have already supplied the brand. A buyer's category question leaves the system to produce the name itself. He would not treat those as interchangeable evidence of visibility, and the benchmark provides no basis for doing so.
Even opening the model doesn't settle it
A third paper, "From Parameters to Answers," examines computation inside the model directly. Researchers asked country-continent questions, then estimated internal signals associated with the country and its continent, removing or reversing parts of those signals while keeping the model's weights fixed. Their conclusion depends on which signal is measured and how it is changed. It provides no universal diagram of how a model fetches a fact from memory.
"A slide labelling your problem a 'recall failure' would need evidence of its own," Dias writes. "Adding technical vocabulary to the slide doesn't supply the missing experiment."
None of the three studies measures the same thing. MemToC tests responses to conflicting tool evidence; Empty Shelves compares strongly cued reproduction with reliable answering; From Parameters intervenes on internal activations. Combining them into a tidy funnel would create a model none of the papers tested.
The chart can still be right
Dias is careful not to dismiss measurement itself. A company may reasonably care whether buyers encounter its name, regardless of the internal mechanism. A carefully defined sample of answers can describe an outcome worth watching, and you don't need to locate a fact inside a model to count a mention.
He tested the limits of mention counting on himself. Two months ago he announced on LinkedIn that he was the "world's most renowned AI visibility expert" — and people searching that phrase still find his original post cited in Google AI Overviews. The answer can explicitly describe the title as a joke and still name him and cite the post. A counter recording only whether his name appeared would tick that answer just as happily as an outright endorsement.
Edward Sturm asked Dias whether the experiment would have worked without his 20-plus years of experience in SEO and information retrieval. Dias told him probably not — but notes the result alone doesn't tell us what part that experience played.
What practitioners should ask
Suppose a brand appears in fewer answers this month. Repeated sampling might show the difference exceeds ordinary variation under the tested conditions. That would establish a change in the measured outcome while leaving its cause open: the model changed, the supplied sources changed, or the questions differed. "We appeared less often" doesn't tell you which to pursue.
Get the diagnosis wrong, and a competent team can spend weeks on work that never addresses the failure. A claim that content is inadequate will send budget toward more content work. If the proposed problem is that the model hasn't learned the brand, the conversation turns to training data.
Dias allows room for action without complete explanation. A team can have good reasons to test an intervention early, and a controlled improvement would give the work a defensible commercial basis even if the mechanism remained partly unclear. "'We have a hypothesis worth testing' is a perfectly respectable starting point for a proposal," he writes.
His closing challenge to vendors: "If you're selling me a remedy for that red cell, what evidence tells you which problem I have?"
Source: Search Engine Journal (https://www.searchenginejournal.com/it-was-there-a-minute-ago/589500/) / Source: arxiv.org (https://arxiv.org/abs/2608.26295)
Source: Search Engine Journal; Source: arxiv.org


