· 8 min read

On September 15, Cloudflare Starts Blocking AI Training Crawlers by Default. If You Run a Content Site, This Is Your New Distribution Question.

On September 15, Cloudflare Starts Blocking AI Training Crawlers by Default. If You Run a Content Site, This Is Your New Distribution Question.

On July 1, Cloudflare rolled out AI Crawl Control to every customer, including the Free plan. It sorts AI bots into three buckets (Search, Agent, and Training) and lets you allow or block each one separately. The headline change lands September 15, 2026: Training and Agent crawlers get blocked by default on pages that display ads, while Search stays allowed. And it doesn't only apply to brand-new sites: it hits new customers, new sites added by existing customers, and existing Free-tier customers who haven't changed their settings. Cloudflare is also evolving its "Pay Per Crawl" marketplace into "Pay Per Use," where publishers get paid when their content creates value in an AI product, not just when a bot fetches it.

If you run a blog, a docs site, a newsletter archive, or anything that lives on written content, this is not a publisher-versus-OpenAI story you can watch from the sidelines. It's a config decision that's about to have a default, and the default flips in September.

Three buckets, three different questions

The useful thing Cloudflare did here is stop treating "AI bots" as one category, because they were never one thing and pretending otherwise forced a dumb all-or-nothing choice.

Search crawlers are the ones that index your content so it can show up when someone asks an AI assistant a question and gets your page cited back. That's discovery. For most solo operators, that's traffic you want.

Training crawlers scrape your content to fold it into the next model's weights. You get nothing back: no link, no citation, no visit. Your words become part of a model that may later answer the exact question your post answers, without ever sending anyone to you.

Agent crawlers sit in between: an AI agent fetching your page in real time to complete a task for a user right now. Whether you want that depends on whether the agent is acting like a visitor or like a scraper with extra steps.

Before this, blocking "AI" meant blocking all three, which meant giving up discovery to protect against training. Now they're separate toggles, available even on Free. That's a genuine improvement. It also means you have a decision to make that you didn't have to make before.

The default is the part that bites

Here's the trap. Starting September 15, the default becomes: Training and Agent blocked on ad-bearing pages, Search allowed. And when you block Training, multi-purpose crawlers get caught in the net: Googlebot, Bingbot, and Applebot can get blocked as a side effect, because a bot that crawls for both search and training can't be half-allowed.

The part that catches people: this isn't only for new sites. Existing Free-tier customers who haven't touched their settings inherit the new default too. If you're a solo operator who set up a domain on the free plan, flipped on Cloudflare, and moved on, you are exactly the person who gets the default applied without reading the notice, and then wonders three weeks later why indexing looks off. You can opt out anytime before September 15 in your Security settings; you just have to know to. Big publishers have someone whose whole job is watching for this. You have a to-do list with forty other things on it.

The specific risk isn't "Cloudflare will hurt you." It's that a default designed for ad-supported publishers who want to fight training scrapers gets applied to your situation, which might be completely different: you might want every bit of discovery you can get and not care at all about training-set purity.

Pay Per Use is real, but be honest about scale

The other headline is monetization. Cloudflare's Pay Per Use extends the earlier Pay Per Crawl idea: instead of only charging bots to fetch your pages, you can get paid when your content actually creates value in an AI product: when it shows up in an answer, when a premium page gets accessed. Launch partners include You.com and Ceramic.ai.

This is genuinely interesting and I'd keep an eye on it. But I'm going to be straight about who it's for right now. Pay Per Use makes real money when you have real traffic and content an AI product wants badly enough to pay for. For a large publisher, that's a new revenue line. For a solo operator with a few hundred posts and modest traffic, it's mostly theoretical: the per-use payouts on a small site round to coffee money, and the setup time isn't free. File it under "watch this become real," not "quit your day job." The infrastructure existing at all is the story; the check clearing for a small site is not, yet.

What I'd actually do

Don't inherit the September 15 default. Go into Cloudflare, find AI Crawl Control, and make the three decisions deliberately (Search, Agent, Training) based on what your site is actually for.

For most solo-operator content sites, the honest answer is: keep Search allowed, because discovery in AI answers is where a lot of new traffic is going to come from, and pretending you can opt out of the AI-mediated web and still get found is wishful thinking. On Training, decide whether you actually care. If your content is your product and you'd rather not seed a competitor's model, block it, but do it knowing you might catch a general-purpose crawler in the process, and check your search indexing afterward. On Agent, block it if you're seeing agent traffic that looks like scraping and want the option, allow it if real users' agents are fetching your pages to help them.

Then screenshot your settings, the same way you'd screenshot any config that has a default that changes under you. Six months from now when your traffic shifts, you want to know what you chose versus what the platform chose for you.

Here's the honest counter-take, and it's the one I keep landing on for my own sites: the instinct to wall off your content from AI is emotionally satisfying and often strategically wrong for a small operator. Your problem is not that too many people know about you. It's that not enough do. For most of us, being findable in AI answers is worth more than the principle of keeping your words out of a training set you were never going to get paid for anyway. Make the choice on purpose, but make it about distribution, which is your actual constraint, not about a fight the big publishers are having on your behalf. I run this blog on a content stack, and I'm keeping Search wide open. The training toggle is the one I'm actually thinking about.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts