Cloudflare Blocks Mixed-Use AI Crawlers on September 15. Here's the Fix Before Your Ad Revenue or Your AI Visibility Takes the Hit.
In three days, on September 15, 2026, Cloudflare flips a default setting that decides who gets to read a huge slice of the ad-supported web. Any crawler Cloudflare classifies as mixed-use, meaning it blends search indexing with AI training or agent fetching under one user agent, gets blocked by default from pages that run ads, unless the site owner explicitly says otherwise. AI crawlers already make up roughly 52% of crawler requests as of June 2026, up from about 22% in spring 2025, according to Cloudflare's own numbers. I sat down this week and actually checked my Cloudflare settings for the first time in months. Most people running a Cloudflare-fronted site haven't, and the window to do it before the default changes under them is closing.
what a "mixed-use crawler" actually is
The problem Cloudflare is solving is that a single bot often does two jobs that publishers feel very differently about. GPTBot, PerplexityBot, and similar crawlers can show up under one identity whether they're indexing your page so a chatbot can cite it in an answer, or hoovering it up as training data with no attribution and no traffic back to you. Search indexing is the deal publishers have accepted for two decades: let Google in, get referral traffic out. Training and pure agent fetching break that deal, since there's often no click, no citation, and no way to tell which purpose a given request served. Cloudflare's fix is to force crawlers to declare intent, search versus agent versus training, rather than hide behind one blended identity. If a crawler won't separate itself out, the new default treats it as a liability on any page carrying ads.
Worth noting: this default change doesn't sweep every Cloudflare account overnight. It applies to new Cloudflare customers, to new sites set up by existing customers, and to all existing free-tier customers. If you're on a paid plan with an established site, your current settings likely hold until you touch them, which makes it easy to assume you're covered when you might not be.
why Cloudflare is forcing the split now
The stated reason isn't just fairness, it's waste. Cloudflare has reported that more than half of the crawl traffic it sees from AI bots is spent re-fetching pages that haven't changed since the last visit. That's bandwidth and origin-server compute spent for zero new information, on top of the compensation question. Cloudflare's CEO framed the broader shift as web traffic itself changing character: bot traffic recently passed human traffic online for the first time, years earlier than expected. Whether or not you buy that framing as the full story, the practical upshot for a site operator is the same either way: a growing share of your server load is coming from something other than a person deciding to read your work, and Cloudflare's answer is to make you choose, explicitly, whether that's a trade you want to keep making for free.
who this actually affects
If your site sits behind Cloudflare and runs any ad unit, you're in scope by definition once the new default goes live on qualifying accounts. But the group that should pay closer attention is broader than "sites with ads." If any meaningful chunk of your traffic comes from being cited or surfaced inside an AI answer engine rather than from search rankings or direct visits, blocking the wrong crawler trades away discovery you can't easily get back. That includes a personal blog like this one. I don't run heavy monetization here, but I do get a steady trickle of visits that started as someone asking ChatGPT or Perplexity a question and landing on a specific post afterward. I genuinely don't know what share of my traffic that is relative to anything ad-driven, and writing this post is the first time I've tried to find out.
the pre-deadline checklist
Here's what I actually walked through, in order:
- Pull your current
robots.txtand check whether it's served from your origin or generated by Cloudflare's managed robots.txt feature. You need to know which one wins before you change anything. - Open Cloudflare's bot management and AI Crawl Control settings and look at the per-crawler list it already shows you, request volume by bot, not just a blanket on/off switch.
- Go crawler by crawler for the ones that actually show up in your logs, commonly GPTBot, ClaudeBot, PerplexityBot, and Google-Extended, and decide: allow, block, or route through Cloudflare's Pay Per Crawl marketplace (now evolving into a usage-based "Pay Per Use" model, currently live with Ceramic.ai and You.com as the first partners).
- If a bot has separate tokens for search versus training (some do, some don't yet), treat them as separate decisions. Blocking the training identity doesn't necessarily block the search one, and vice versa.
- Confirm which of your pages actually carry ads or monetization, since that's the trigger condition for the new default, and check that any manual overrides you set actually persist past September 15 rather than getting reset by the policy change.
- Put a date in your calendar two to four weeks out to look at both your ad numbers and your referral sources from AI platforms, so the decision you make this week isn't the only one you ever make.
# illustrative example only, check your own logs before copying this User-agent: GPTBot Disallow: / User-agent: PerplexityBot Allow: / User-agent: * Allow: /
the honest take
Blocking every AI crawler protects whatever ad revenue you currently have, today, with certainty. Allowing them keeps you eligible for whatever AI-driven discovery turns into, later, with no certainty at all, since Pay Per Use only pays out through partners who've actually signed up, and most AI companies haven't. That's a real tradeoff, not a solved problem, and the honest answer depends on which side of it your traffic actually leans on. Most solo site owners, myself included until this week, have never measured that split. We assume ads matter because they're the line item we can see, and we assume AI citations matter because everyone's talking about them, without ever pulling the actual numbers.
What I'd actually do: don't block everything by reflex just because the deadline is close. Check your analytics for referral traffic that looks like it came from an AI answer engine, weigh that against what you'd lose in ad impressions if you left training crawlers unblocked, and set your per-crawler rules based on that, not on the calendar. Where I could be wrong: if your ad revenue is trivial and your AI referral traffic is also trivial, none of this changes much for you either way, and the bigger risk is just spending an afternoon on settings that don't move any number you care about. For a site my size, that's a real possibility, and I won't know until I check back in a month.
Author
Lukas
@lukcombinatorSources
- Cloudflare gives AI crawlers a September deadline: pay publishers or get blocked
- Cloudflare's new policy pushes AI companies to pay for publishers' content
- Content Independence Day, one year on: building the business model for the agentic Internet
- Control content use for AI training with Cloudflare's managed robots.txt and blocking for monetized content