AI Citation Study: Source Position Matters Less Than Raw Data Suggests
A GPT-5.4 preprint test found top results cited 85.1% vs 42.8% for fifth position, yet swapping source order produced a zero-point effect in a follow-up set.
Updated

The brief
- Top-position pages were cited 85.1% of the time versus 42.8% for fifth position, but swapping source order within matched pairs yielded a 0.0-point estimate (-5.4 to +5.4, 95% CI) in a 56-pair set.
- Structured rewrites earned 0.50 more citation markers per answer (95% CI 0.20–0.84), though the pre-planned test of being cited at all (+4.5 points) was inconclusive.
- Reruns with identical inputs flipped the cite/no-cite decision in 15% of cases; the authors estimate ~45% of single-run variation comes from model randomness.
A new preprint study tested whether changing one part of a source alters AI search citations, holding everything else constant. The raw numbers look dramatic: the top result in a search call got cited about twice as often as the fifth. But when researchers Sriram Selvam and Anneswa Ghosh reversed the order of matched sources, the effect shrank sharply — and measured zero in a follow-up test.
Posted to arXiv on September 14, the paper is not peer-reviewed. It covers a single GPT-5.4 search agent using Exa as its search provider, with offline replayed conversations and no live webpage edits.
How the test worked
The researchers prompted the GPT-5.4 agent to answer 130 common questions through independent web searches. They recorded every message and search result from the 129 questions it addressed. From these transcripts, they selected pairs of pages that appeared in the same search outcomes and both passed screening as supporting the same fact. The goal: isolate situations where either page could be fairly cited, so any credit differences reflected only how the model split recognition between two sources confirming the same fact.
That left 113 pairs. A later blinded human check confirmed 103 as genuine matches. The researchers replayed each saved conversation four ways — one page above or below the other, and its text shown either as plain paragraphs or rewritten with headings, lists or a table. Only the final answer was regenerated.
Both text versions came from AI rewrites of the original page. Grok 4.3 created nearly all of them, with GPT-5.4 as a fallback for one pair, and a separate Grok review verified the facts matched. Because wording varies between the two versions, the authors note the test compares two rewrites and does not isolate formatting differences alone.
The raw position gap was larger than swap effects
In the initial position of a search call, pages were cited 85.1% of the time in saved transcripts, versus 42.8% for pages in the fifth position — a gap of 42.3 percentage points. Here, "position" means the order of the five Exa results returned in a single search, not Google rank or live-web placement.
The study points out that search providers usually place more relevant pages at the top, so the raw difference blends position with page quality. When the researchers moved the same page higher within its pair, the chance it was cited at all rose by 7.9 percentage points — not statistically significant after accounting for multiple tests.
A separate set of 56 pairs, where only the order was switched, showed an estimate of 0.0 points, with a 95% confidence interval from -5.4 to +5.4.
The study explains that the raw gap and the swap results measure different things. Position did influence citations in some cases, it concludes, but averages from raw position data are not reliable.
Structured rewrites got more credit, not clearer entry
Pages rewritten with headings and lists received an average of 0.50 more citation markers per answer than the same pages in plain-paragraph form, with a 95% confidence interval from 0.20 to 0.84. The test's answers were heavily cited, with a median of 29 markers across six documents.
The total number of citations per answer did not rise, and the count on the other page barely changed. The authors read this as credit concentrating on the rewritten page.
The researchers' pre-planned primary test — whether a page got cited at all — found structured text increased that likelihood by 4.5 percentage points, with a 95% interval from -1.4 to +10.4. The paper calls this inconclusive and notes the study could reliably detect only effects of roughly 8.5 points or larger.
A stricter comparison, where every word stayed identical but the layout changed to one sentence per list row, boosted citation rates across all 113 pairs. Repeating the test on a subset reversed the effect.
In the discussion section, the authors wrote: "This is an attribution-sensitivity warning, not an optimization tactic."
Reruns changed citation outcomes
The researchers retested 120 responses with identical inputs. The decision to cite or not cite the target page changed in 15% of cases — roughly one in seven. The average count effect held steady across reruns. They estimate about 45% of the variation in a single run's effect comes from model randomness.
The authors recommend rerunning citation tests multiple times and reporting consistency across runs. Separately, SparkToro reported in January that ChatGPT and Google's AI Overviews each produced the same brand list less than 1% of the time when given the same prompt repeatedly.
Why this matters
The raw position gap in this test far exceeded the average effect observed when researchers actually swapped source order. An Ahrefs report from May showed pages cited by AI were about three times more likely to include JSON-LD schema, yet adding schema did not clearly increase citations.
That raises a question for marketers: was a correlation in a vendor report — or in your own tracking — ever tested by changing the variable itself? A single answer is a weak basis for labeling a citation as won or lost.
The study cannot confirm whether reformatting a live page boosts citations, since rewrites applied only to text already retrieved, excluding crawling, retrieval and ranking processes.
Looking ahead
The researchers re-ran the saved searches on Grok 4.3 and found the structured rewrites leaned the same way. However, fewer than half of Grok's first replies followed the correct citation format. They recommend further research testing each scenario multiple times, across different search providers and models, tracking both citation frequency and whether a page gets cited at all.
Based on Search Engine Journal; arxiv.org
Filed under ai-search, citations, seo, study, gpt
Tom Whitfield
Show full bio
Staff writer covering consumer brands and retail at Marketing Herald.
More from the wire
- Chris Green Maps the Data Sources Feeding AI Search Results
- Field Experiment: Google AI Mode Cuts Clicks and Degrades UX
- AI Visibility: How to Map the Sources That Shape Answers in Your Industry
- SEO Hack: Checking If Your Pages Feed AI Search Retrieval Pipelines
- LinkedIn as an AEO Channel: A Solopreneur's Three-Week Experiment