SEO & Search
Cloudflare Will Write Your Robots.txt, And It Has A Point
Cloudflare's Bot Preference Sync will auto-generate robots.txt from three dashboard settings and block AI crawlers failing its four disclosure conditions. Defaults change for new domains on September 15.

Cloudflare has announced Bot Preference Sync, a product that writes your robots.txt for you. Whatever bot policy you set in Cloudflare's dashboard gets converted into robots.txt entries and prepended to your file between # BEGIN Cloudflare Bot Preference Sync and # END Cloudflare Bot Preference Sync markers, with your existing content preserved underneath. Cloudflare says the feature will run from the free tier up and will be on by default for new customers.
As of September 13, there was no entry for Bot Preference Sync in Cloudflare's bots changelog — which still ends at July 1 — no mention in Cloudflare's bots documentation, and no generated block on the original author's robots.txt. The product exists so far only as announced.
Why the gap between file and enforcement matters
Cloudflare's announcement states that "when your stated preferences and your enforced rules disagree, some crawlers treat it as a basis to disregard your preferences or try to bypass your enforced rules." Cloudflare does not name which crawlers, how many, or how it knows. The company sits in front of a large share of web traffic and is the one party positioned to name them, and it names none.
The gap is easy to create because robots.txt and edge enforcement live in different places. The file is text written once, probably long ago. Enforcement is a dashboard changed at some point since. Nothing keeps them aligned. The author of the source piece found their own robots.txt had welcomed Bytespider by name for months after they would have chosen otherwise — caught only while writing a reference page on robots.txt enforcement, published August 20. They fixed the file by hand between August 20 and August 26.
Three categories, not per-crawler control
Bot Preference Sync generates robots.txt from three settings under Security Settings and Configure AI bot policies: Search, Agent, and Training. Each offers: block on all pages, block only on pages with ads, or allow. The Training setting differs — the announcement describes it as a Disallow option that writes a no-training line into the file. Cloudflare's tracked bot list decides which crawlers fall into which category. You cannot exclude an individual bot from the sync. For finer control, Cloudflare's stated answer is to turn the sync off and maintain the file yourself.
That is a real limitation. The source author allows OpenAI's GPTBot, Anthropic's crawler and PerplexityBot, and blocks Bytespider and meta-externalagent. Every one of those companies trains models. The policy is a business decision made per company: the first three put pages in front of people asking assistants questions; the other two take and return nothing. Set Training to disallow and the file tells OpenAI not to train on content OpenAI is welcome to train on. Set Training to allow and nothing separates Meta and ByteDance from anyone else, while the edge returns 403 to both. No setting describes what the site actually does.
Four disclosure conditions, with blocking attached
Setting Training to disallow blocks every AI crawler Cloudflare judges opaque. Cloudflare has published four conditions a mixed-use crawler must meet to avoid that:
- It "must respect, via any mechanism, a 'no training' preference in robots.txt"
- "They give site owners a way to opt out of AI summaries"
- "They provide URL-level visibility into which pages were made available for training, as well metrics on search results, so you can see how your content was used for search and for training"
- "They can show publicly that Disallowing Training does not hurt your traditional search results"
No company is named. Conditions two and four describe Google. Google already meets condition four: its crawler documentation, last updated July 14, 2026, says Google-Extended controls whether crawled content trains Gemini models and "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." Condition two has no Google answer. Google's AI features documentation lists controls — "nosnippet, data-nosnippet, max-snippet, or noindex" — every one of which limits Search everywhere. Its rule that a page must be indexed and snippet-eligible to appear as a supporting link in an AI Overview means one switch governs both. There is no setting that exits AI Overviews while leaving ordinary snippets alone.
Microsoft answered condition two in September 2023: content tagged NOARCHIVE "will not be included in Bing Chat answers" and "will still appear in our search results." Cloudflare's July 2026 post on AI traffic options names BingBot alongside Googlebot as the mixed-use crawlers these conditions govern.
The terms are reasonable, but a vendor wrote them, a vendor's network enforces them, and the websites doing the blocking mostly clicked one toggle and never saw the conditions.
Defaults changing September 15
On September 15, Cloudflare changes defaults for all new domains: Training and Agent blocked on pages that display ads, Search allowed. Until then, a new customer who sets nothing gets no blocks — Cloudflare says "the starting point will not add any blocks on your behalf." Selecting "I monetize from pages with ads on this domain" during onboarding sets Training to Disallow automatically. A question about your business model becomes a published position on AI training in your robots.txt.
Cloudflare already writes its analytics script into free-plan websites unless you opt out, as its blog announced in September 2025 and an August 2026 Hacker News thread rediscovered. The same on-by-default pattern now reaches robots.txt, a file people open even less often than a settings page.
What robots.txt actually stops
A robots.txt rule stops a crawler that chooses to be stopped. The author measured the other kind on their own site: the biggest so-called AI crawler in their logs was hunting for credentials — requesting /.env and SSH keys under a nonprofit research archive's name. Enforcement remains at the edge. The file documents intent, which matters later in a dispute, not when a request arrives.
Three checks take about 10 minutes: read your robots.txt, including parts written in 2023; compare it against your Cloudflare AI bot policies; and decide whether the three categories can express your policy. If your policy is "open to everyone" or "closed to training," Bot Preference Sync saves a manual job. If your policy is per company, turn the sync off and maintain the file yourself. And when the sync reaches your site, read what it wrote — a policy file you have not read is a statement someone else is making on your behalf.
Source: Search Engine Journal (https://www.searchenginejournal.com/cloudflare-will-write-your-robots-txt-and-it-has-a-point/589262/) / Source: nohacks.co (https://nohacks.co/blog/what-courts-say-about-blocking-ai-bots)
Source: Search Engine Journal; Source: nohacks.co



