MH-4228SEO & Search
SEO Developer Tests Gemini Nano in Chrome, Settles on Three-Layer AI Architecture
Chris Green tested Chrome's built-in Gemini Nano for technical SEO tasks. Nano failed at judgment calls, so he built a three-layer architecture: exact code, local interpretation, frontier models on demand.
Wire notes
- Chris Green benchmarked Chrome's built-in Gemini Nano against Gemini Flash and ChatGPT Luna and found Nano unreliable at final judgment calls on technical SEO evidence.
- His resulting extension uses a three-layer architecture: deterministic code for exact tasks, Nano for light interpretation, and larger API models (Gemini, OpenAI) for genuine reasoning.
- Green argues frontier models are unnecessary for many tasks and that current compute costs are unsustainable, predicting browsers and OSes will ship better local models over time.

Gemini Nano, the tiny model Google ships inside Chrome, cannot reliably make the final judgment calls in technical SEO work — but it can still take a meaningful amount of work off the user, according to Chris Green, the SEO developer behind the Exactly Matchy tool.
Green, writing on Chris Green SEO and syndicated by Search Engine Journal, ran a series of experiments to answer what he calls "a more interesting question": How much useful work can we move closer to the user?
The context matters for marketers building AI-assisted tooling. Much of the current AI conversation assumes that "useful" AI means getting an agent to do the whole thing for you, Green argues. In SEO, GEO and AEO work, there are tasks where a degree of interpretation is useful, but sending every request to a large remote model is not needed or even the best option.
"If I want to extract & deduplicate a URL list collection from XML sitemaps, you don't need a frontier model," Green writes. A cheap, predictable script does that job better.
Why local, and why not
Running AI locally means using your own hardware — phone, computer, laptop — to do the compute work without sending it elsewhere. ChatGPT or Claude are simple and often free, but Green lists the drawbacks: they are resource-intensive in terms of data centers and water use, costly and getting more expensive, raise data privacy questions, and put in a point of failure you cannot control.
Local inference, by contrast, requires no API call for every minor task, keeps data on-device, can be fast if model and session startup are handled well, and lets tools you build keep working without depending on an external AI service.
When Green built Exactly Matchy — a tool that helps you understand if your content is retrievable by AI systems — he deliberately avoided APIs, credit cards and what he calls "general faffery," leveraging Chrome's built-in version of Gemini Nano, which downloads when needed.
Small does not mean simple
The experiments exposed a hard limit. Green tested whether Nano could analyze the differences between raw HTML and the rendered DOM — evidence such as an anchor tag with the same anchor but a different destination, the same destination but different anchor text, a broken destination in the initial HTML that becomes working after rendering, or two different URLs that resolve to the same final destination.
Experienced SEOs would find that data crucial to understanding whether a raw/rendered difference is a real problem. Green's verdict on Nano: "Sadly not… This kind of decision-making is not an easy reasoning problem; it seems easy, but it isn't simple."
The model still has to understand what the evidence means, respect the facts, combine multiple signals and avoid inventing information. Nano is an intentionally small, fast model, quantized to fit in Chrome without slowing things down — and that design is not ideal for judgment work.
"In testing, Nano was useful at some tasks but unreliable at making the final judgment you could really trust," Green writes. "A stronger API (ChatGPT or Gemini) model handled the same evidence considerably better."
Three layers, one pipeline
The new extension Green built settled into a three-layer architecture.
Code handles things that should be exact. Fetching URLs, comparing HTML, checking HTTP responses, matching elements, identifying canonical relationships and detecting destination changes do not need probabilistic reasoning. "If anything, asking an LLM to answer these questions is risky," Green writes.
A small local model handles light interpretation and communication. Once the facts are established, Nano can turn an ugly bundle of evidence into something a human can use quickly. "Friction kills progress more than almost anything else," Green says, arguing that presenting a readable passage rather than blocks of JSON is highly valuable. He considers it a strength that Nano doesn't have to make the decision.
A larger model is available when actual judgment is needed. For complex, ambiguous or semantic reasoning, the same structured evidence goes to Gemini, OpenAI or another capable model. The pipeline does not need to change — only the model does.
Benchmarks miss the point
Green benchmarked Nano's reasoning against Gemini Flash and ChatGPT Luna, treating the models equally. That framing, he argues, is where thinking about different AI models goes wrong: "They do not need to replace frontier models to be useful! You certainly don't need a frontier model for everything either!"
He also flags cost as a strategic concern: "To be brutally honest, I don't think compute costs as they are today are sustainable, so maybe we shouldn't overly rely on it."
A side effect of supporting a small model pushed Green to improve the deterministic code, which made the evidence better, focused Nano's responsibilities and — as a by-product — made the stronger models perform better too.
Chrome's current local model will not be its last, Green predicts, and browsers, operating systems, laptops and phones will ship improved models, quantization methods and hardware. Applications designed around a replaceable local model can absorb those improvements without a redesign.
The opportunity, Green concludes, is not to recreate ChatGPT or Claude locally. It is to build software where "exact computation happens in code, lightweight intelligence happens locally, and expensive intelligence is called only when it is truly needed" — a direction he calls much more sensible than treating every problem as an excuse to spin up the largest model available.
via chrisgreenseo.substack.com (Original)
More from James Calloway
Show full bio
Market editor covering media and advertising at Marketing Herald.
60 articles
More on the wire
- Google: AI Search Demands Location-Level Data, Not Brand-Level Presence
- Field Experiment: Google AI Mode Cuts Clicks and Degrades UX
- Google, Cloudflare, Microsoft: How AI Content Payment Models Differ
- Tinuiti CEO: AI Agents Fail Without Clean Data Foundations
- Chris Green Maps the Data Sources Feeding AI Search Results