SEO & Search
Chris Green Maps the Data Sources Feeding AI Search Results
Chris Green tiers the data sources behind AI search, from confirmed Google Search grounding and Yelp bookings to a $60m Reddit training deal marketers should treat as unstable.

Search strategist Chris Green has published a tiered reference of the data sources that AI search tools actually pull from — and only a fraction of them are things most marketers actively manage.
Green, writing on Search Engine Journal, argues that the industry suffers from "search source myopia": a short-sightedness around which sources matter for day-to-day work. As AI chatbots and tools like Google's AI Overviews and Microsoft's Copilot draw from an increasingly diverse range of data sources, he writes, "our job gets a lot harder. And more interesting."
His answer is a four-tier classification. Tier 1 covers confirmed, current sources used for retrieval-augmented generation, grounding or actions. Tier 2 covers confirmed, current sources tied to training or licensing deals. Tier 3 covers confirmed historical pretraining corpora. Tier 4 covers sources with strong evidence but no confirmation. Green advises approaching Tier 1 "with the most interest — as they'll more-than-likely be worthwhile," while noting Tier 2 and 3 may be harder to act on.
What Tier 1 confirms
The confirmed, live grounding sources include Google Search, which connects Gemini to real-time web content with inline citations to source URLs, and Bing, which Microsoft documents as enhancing Copilot responses.
In local search, the Tier 1 list runs long. Google Maps provides geospatial grounding through Google Cloud's Grounding API. Google Business Profile feeds Google's local surfaces. Yelp licenses reviews, photos and business information to OpenAI for real-time local recommendations — and the integration goes beyond retrieval: ChatGPT users can book a table, join a waitlist, or use Request a Quote to contact providers in-chat. Yelp's 10-Q confirms the deal is live.
Wikipedia and Wikimedia sit in Tier 1 as both a pretraining corpus and a live reference corpus, with unusually clear licensing — principally CC BY-SA, with attribution and share-alike obligations.
The Google–Reddit deal also earns Tier 1 status on the grounding side: it gave Google access to the Reddit Data API for "real-time, structured, unique content" and allows Reddit content to appear across Google products. The training side of the same deal, reported at roughly $60 million per year, sits in Tier 2 — with a warning. Green notes Reddit is reportedly weighing whether to renew, and marketers should "treat as unstable."
Commerce and travel are converging
The most commercially consequential entries involve transactional data. Google Merchant Center feed data underpins Google's shopping surfaces. On the OpenAI side, merchants share a secure, regularly refreshed CSV or JSON feed of identifiers, descriptions, pricing, inventory, media and fulfilment so ChatGPT can surface products accurately — with refreshes accepted as often as every 15 minutes.
Travel follows the same pattern. When Gemini or AI Mode show hotel options with real-time prices, that data comes from Google Hotels feeds. In August 2026, Google added hotel booking inside AI Mode completed with Google Pay — grounding plus actions in a single flow. Booking Holdings and IHG are reported participants in Google's agentic booking pilot, though Green classifies partner feeds as Tier 4 because no documented feed spec exists.
Licensing deals dominate Tier 2
OpenAI's publisher partnerships anchor the second tier: deals with the Financial Times, Axel Springer, AP and News Corp are all confirmed, with terms differing per partner on training versus grounding versus attribution. Axel Springer's deal includes otherwise paywalled material in answers. AP licensed part of its text archive.
Developer sources also land in Tier 2. LLaMA used GitHub's public BigQuery dataset, restricted to Apache, BSD and MIT projects — and Green cautions that "public visibility is not an open licence." Stack Overflow appears in licensing-deal mapping alongside Reddit and Shutterstock as a data platform powering multiple buyers.
Tier 3 and the inferences
Historical pretraining corpora fill Tier 3. GPT-3 used filtered Common Crawl as roughly 60% of its sampling mixture; LLaMA 1 reported 67%. C4, a cleaned derivative of Common Crawl, accounted for 15% of LLaMA's pretraining mixture.
Tier 4 contains the educated guesses: OpenStreetMap, Foursquare, Tripadvisor, Wikidata, social platforms and forums. Green deliberately leaves Foursquare there because the OpenAI local deal is Yelp's, and no equivalent evidence exists. He flags npm and PyPI registries as "the weakest entry in the table," with no disclosed agreement or documented retrieval use.
The practical play
Green acknowledges the table "may date badly — an occupational hazard of 'AI Search' at the moment." His recommendation is to look ahead to data sources AI providers will want next, and to adapt by market: where listed sources have less traction — Yelp, he notes, isn't huge in the UK — marketers should investigate other big regional players even without a confirmed relationship.
The most meaningful research, he writes, is studying which of these sources are shaping AI results today, by examining generated results for queries customers actually search — and looking for the gaps.
Source: Search Engine Journal (https://www.searchenginejournal.com/which-data-sources-should-you-care-about-ai-search/590086/) / Source: chrisgreenseo.substack.com (https://chrisgreenseo.substack.com/p/which-data-sources-should-you-care)
Source: Search Engine Journal; Source: chrisgreenseo.substack.com




