webscrape.dev

Cloudflare put ai-train=no in 3.8 million robots.txt files. Your price scraper can ignore it. Your training set can't.

Cloudflare's Content Signals policy put ai-train=no into 3.8 million robots.txt files by default. What that advisory line means for teams collecting web data.

Nathan Kessler

Written by Nathan Kessler

Last updated: 7 min read

In late September 2025, a new line appeared in more than 3.8 million robots.txt files at once. Almost none of the site owners typed it. Cloudflare added it for them, as a default, when it launched its Content Signals Policy and switched it on across the domains that use its managed robots.txt. If you collect web data at scale, this is the shape of the thing you now meet on a large slice of the web: a content signals robots.txt directive that reads ai-train=no, sitting three lines under the Disallow rules your crawler already obeys.

The first instinct is to treat it as another block to route around. It is not a block. What it actually is, and which of your pipelines it speaks to, matters more than any workaround, because the line changes what your legal team can say later more than it changes what your crawler can fetch today.

What is the Content Signals policy in robots.txt?

The Content Signals Policy is a robots.txt extension that states how a site's content may be used after it has been fetched. It adds a Content-signal: line carrying three tokens, each set to yes or no: search (build a search index and show links and snippets), ai-input (use the content as input to a model at run time, such as RAG, grounding, or AI Overviews), and ai-train (train or fine-tune a model). A token that is absent means no preference. It is a stated preference, not a technical block, and it is not enforceable by itself.

Cloudflare launched the policy on September 24, 2025, and turned it on by default for more than 3.8 million domains on its managed robots.txt. The default string it inserts is Content-signal: search=yes, ai-train=no, with ai-input deliberately left unset so the site owner can decide. So the common case you will hit is a site that says, in machine-readable form, "indexing is fine, training is not, and I have not said anything about inference."

It marks use, not access

Every robots.txt rule you have handled until now is about access. Disallow says do not fetch this path. Crawl-delay says fetch it slower. The Content Signals line is different in kind. It says nothing about whether you may request the URL. It governs what you do with the bytes after they are in your pipeline.

That is why a fully robots-compliant fetch can still cross an ai-train=no. Your crawler reads the file, honors the Disallow blocks, and pulls the page it is allowed to pull. Nothing in that exchange touches the Content-signal line, because that line is not an access rule and a standard parser has no reason to act on it. The request looks clean. Whether you then feed the page into a training corpus is a decision made downstream, in a different system, often by a different team. The signal is aimed at that downstream decision, and the fetch layer never treats it as its problem. For the background on why access-layer blocks and this use-layer signal are separate concerns, why scrapers get blocked covers the access side in full.

A CDN set the default, not the site owner

Here is the part that matters most when you weigh how much the signal should bind you. On the large majority of those 3.8 million domains, no human chose ai-train=no. Cloudflare chose it as a platform default and applied it wholesale. The line is present because a CDN inserted it, not because a publisher sat down and reserved a right.

Google's search relations team made the same point bluntly. John Mueller dismissed the directive in mid-2026: "none of the crawlers / llms use the 'content-signal' robots.txt directives. It was made up by a CDN," he wrote, adding that it "has no effects whatsoever for any crawler or llm" and "just adds bloat & future maintenance to your robots.txt file." He is right about the mechanics. Nothing about the line forces any crawler to behave differently, and the major crawlers do not read it.

But "no technical effect" and "no significance" are not the same claim. A machine-set default is weak evidence of any specific publisher's intent. A machine-readable reservation, once it exists in a documented format, is still a reservation you can no longer say you never saw. Those two facts pull in opposite directions, and which one wins depends on what you are collecting for.

Does it bind your crawler?

Split the question two ways: by token, and by your use.

By token, the default only asserts search=yes and ai-train=no. If you run a nightly price scrape, a SERP collector, or a search-style index, the signal that names your activity is search, and its default value is yes. The ai-train=no on those same millions of domains is not addressed to you. You are not training a model on the retailer's product pages; you are reading a price. Reading the line correctly means noticing that it does not speak to most conventional collection at all.

The ai-input gap is worth its own note, because it is the token most likely to matter next and the one the default leaves blank. On the 3.8 million sites that carry only Cloudflare's default, ai-input has no value, so a retrieval pipeline that feeds pages into a model at query time gets no preference either way. Only an ai-input=no that an owner set on purpose speaks to inference, and a deliberately set token carries the intent that a wholesale default does not. If you build RAG systems, the sites to watch are the ones that changed the default, not the millions that kept it.

By use, the calculus flips for one group. If you are assembling a corpus to train or fine-tune a model, or building a retrieval set that grounds a model at inference time, then ai-train=no and an explicit ai-input=no are pointed straight at you. This is where the machine-readable part stops being cosmetic. In the EU, machine-readable reservation of rights is the specific trigger that copyright law ties AI-training compliance to, a frame the RSL licensing standard post works through in detail. The short version: for a training or RAG dataset, a machine-readable no is a different kind of fact than a line buried in a terms-of-service page, and "we did not know" gets harder to say with a straight face. For a price monitor, none of it applies.

Where it sits next to the other new signals

The Content Signals line is one of three machine-readable layers that landed on the same file inside a year, and they answer different questions. RSL, the Really Simple Licensing standard, is publisher-authored and can attach a price and license terms to specific uses; it is a licensing offer, not a preference flag. Cloudflare's pay-per-crawl meters declared AI crawlers at the network edge and can return a 402 for payment; it is enforcement, not advice. Content Signals sits between them as the advisory layer. It states a preference and stops there.

Do not confuse that advisory line with Cloudflare's harder move on September 15, 2026, when it set new enforcement defaults in its AI Crawl Control product. For new domains onboarding to Cloudflare, Training and Agent crawlers are now blocked by default on ad-monetized pages, while Search stays allowed. That is a real edge block that returns an error, not a preference you can ignore. The Content-signal line and the Crawl Control block can both sit in front of the same origin and do entirely different things, and a collection team needs to tell them apart before deciding whether a target is even reachable. Standardization is still open: the IETF's AIPREF working group is drafting a common AI-preferences vocabulary, and Cloudflare's token set is an interim extension a ratified standard may later replace.

What a collection team should do

None of this calls for stopping a crawl. It calls for reading a file you already fetch with a little more care, and routing what you find to the right place.

  • Log the Content-signal line. Capture it per target the way you capture status codes. You cannot reason about a signal you never recorded, and if a no ever becomes a question, the record of when it appeared is the artifact you will want. Fold it into the pipeline the way large-scale crawl architecture treats any other per-origin metadata.
  • Classify your own use against the three tokens. Map each pipeline to search, ai-input, or ai-train. That mapping is what tells you whether a given site's signal is aimed at your job or at someone else's.
  • Weight a set token over a default one. An ai-input=no or ai-train=no that an owner changed by hand says more than the value Cloudflare wrote for 3.8 million sites at once. When you have the request logs to tell the two apart, keep the distinction.
  • Treat a reservation as one for the pipelines that train or ground models, and only those. For an ordinary price or SERP scrape, the default signal does not name your use and does not change your posture. For a training or RAG build, it is a reservation worth escalating to whoever owns that risk.
  • Put it in provider and target diligence. When you evaluate a web data provider or a new target list, add "what does this source signal, and does it name our use" to the checklist, next to the questions you already ask about blocking and web scraping coverage.

The line in those 3.8 million files is easy to read as a wall and easy to read as noise. It is neither. It is a use preference, set mostly by default, that reaches your training corpus and skips your price scraper. Knowing which of those you are running is the whole job.

Share:

Tags:

  • #compliance
  • #crawling
  • #legal
  • #ai-training