News Box

SEO & Search

AI Answers Strip the Signals Buyers Used to Judge Evidence

Pew tracked a 1% source-click rate on AI summaries; Wharton experiments show AI users learn less. Marketers must reposition content for confidently underinformed buyers.

Updated

LLMs Are Time Machines That Don’t Tell You How Far You Went via @sejournal, @DuaneForrester
LLMs Are Time Machines That Don’t Tell You How Far You Went via @sejournal, @DuaneForresterhongxing128 / Openverse

The brief

  1. Pew Research Center tracked 900 U.S. adults across 68,879 Google searches in March 2025: users clicked a source cited inside an AI summary on roughly 1% of visits, versus 8% for normal results (15% when no summary appeared).
  2. Wharton professors Shiri Melumad and Jin Ho Yun ran seven experiments with 10,462 participants, published in PNAS Nexus in October 2025, finding AI-summary users knew less, engaged less, and produced sparser, less original advice even with identical facts.
  3. A 2025 Microsoft Research and Carnegie Mellon survey of 319 knowledge workers across 936 real AI uses found greater confidence in AI predicted less critical thinking.

When an AI summary appears in Google results, users click a cited source on roughly 1% of visits. That figure, from Pew Research Center's tracking of 900 U.S. adults across 68,879 Google searches in March 2025, frames a problem Duane Forrester calls the quiet cost of answer engines: the machine hands you a decision and deletes the evidence you would have used to evaluate it.

Forrester, writing on Duane Forrester Decodes, argues that LLMs are time machines that compress research without recording the distance traveled. A library-era question took days. Search compressed it to hours. Answer engines compress it to seconds. The speed is new. The uncertainty underneath is not.

What disappeared, he writes, is what he calls path metadata — the byproduct signals that accumulated during research and told you what an answer was worth. Three books and no journal articles meant the topic was thin. Contradicting sources meant it was contested. A question that took four hours felt different from one that took four minutes, and that difference was information. An answer engine returns a conclusion in the same confident prose whether the evidence behind it was deep or nearly absent.

The research caught up

For years this argument was plausible but unproven. Two Wharton marketing professors, Shiri Melumad and Jin Ho Yun, changed that. Their seven experiments, covering 10,462 people and published in PNAS Nexus in October 2025, had participants learn about ordinary topics — planting a vegetable garden, spotting financial scams — from either an AI summary or standard Google links, then write advice for someone else.

The AI group came away knowing less, even when both groups saw identical facts. They spent less time engaging. The advice they wrote was sparser, less original, and less likely to be adopted by the people who read it.

One finding closed the obvious escape hatch. When the model supplied live web links alongside its answer, participants did not click them. Once the summary arrived, the sources beside it stopped being interesting.

Pew's field data points the same way. When an AI summary appeared, people clicked a normal search result on 8% of visits, against 15% when no summary appeared. They ended the browsing session entirely on 26% of pages with a summary, versus 16% without one. Pew calls this association rather than proven cause, and the study covers one month, U.S. users and Google only. But a controlled experiment and a passive tracking study, using different methods and different populations, reached the same destination.

A third thread predates the technology. In 2015, three Yale researchers ran nine experiments and found that searching the internet inflated how much people believed they knew — even when searches returned nothing at all. Forrester concedes the point against his own thesis: the old journey never taught good judgment. What it left behind was friction, elapsed time and visible variety in sources, and Melumad and Yun show those were doing work. Remove them and the deficit shows up in what people produce.

A 2025 paper from Microsoft Research and Carnegie Mellon, surveying 319 knowledge workers about 936 real uses of AI at work, found that greater confidence in the AI predicted less critical thinking, while greater confidence in one's own ability predicted more. Forrester notes the study is self-reported, correlational, and that one co-authoring institution sells a generative AI product. Dirk Lewandowski's 2026 study of information regret points the same direction from a smaller sample.

The correction that used to happen for free

Under the old model, a wrong answer was survivable. Someone read a bad summary, kept going, landed on the real page and swapped the wrong version for the right one. Nobody designed that repair. It cost nothing, ran millions of times a day, and functioned as what Forrester calls the immune system of the information economy. At a 1% source-click rate, it does not fire.

That asymmetry is the expensive part. The old repair was free and fired in seconds. Whatever replaces it must be paid for: publish the evidence, then wait for it to be crawled, retrieved, weighted or trained on, with no control over the timeline and no confirmation it worked. A misrepresentation that used to be temporary now stays put.

Citation is presence, not traffic

If roughly one visit in a hundred produces a click, treating AI citation as a referral channel values it wrong, Forrester writes. The value is being inside the answer a person acts on. The thing to measure is the answer itself, not the trickle that escapes it.

The inbound lead has changed too. They are not uninformed; they are, in Forrester's phrase, confidently underinformed. A decade of content strategy assumed a staircase — definitional explainers at the top, comparisons in the middle, depth at the bottom. The top of that staircase now happens before anyone reaches you, and the people who arrive are further along the decision journey. But further along and better informed are not the same thing. The visitor carries the confidence of someone who finished the research alongside the depth of someone who read one paragraph.

Beginner content now talks down to them, which ends a visit quickly. Advanced content assumes a vocabulary they can repeat but have not earned, which ends the visit more politely. Marketers miss in both directions at once, using assets built over ten years and a large budget.

Deleting the 101 material is not the answer — it still feeds what the models say, credited or not. Placement is what has to change. The material only the publisher can produce, long positioned as the depth layer, now has to do the job of the front door as well.

The mirror

Forrester closes by pointing the argument at marketers themselves. The competitive analysis, the strategy deck, the recommendation about to go to a client or board — if it came from a model's synthesis, it is the sparser, less original output Melumad and Yun measured, delivered with more certainty than the process earned.

"The time machine works," he writes. "But it drops you somewhere without telling you how far you traveled or what you went past. Not because the answers are wrong. Because you cannot tell."

Based on academic.oup.com

Filed under ai-overviews, search-behavior, answer-engines, content-strategy, click-through-rate

Share this article:

Amara Osei

Show full bio

News editor covering industry trends and analytics at Marketing Herald.

More from Amara Osei →

More from the wire

Previous article