Observed Signal · Jul 1, 2026 · Policy Update · Source: techcrunch · Impact: 4/5 · Sentiment: Positive
Cloudflare blocks mixed-use crawlers on ad pages
Cloudflare announced a policy change that, starting September 15, 2026, will by default block “mixed-use” web crawlers (those that combine search, agent use, and training) from crawling pages that host ads. The new defaults apply to new Cloudflare customers, new sites of existing customers, and all existing free customers unless site owners change their settings. Cloudflare says the change encourages AI companies to separate search crawlers from agents and training crawlers and enables new commercial opportunities for publishers, evolving its Pay Per Crawl marketplace into a Pay Per Use model. Initial partners for publisher payments are Ceramic.ai and You.com. Cloudflare also cited internal data showing over 50% of AI crawler traffic re-fetches unchanged pages, and framed the move as protecting publishers’ IP and bandwidth while reshaping access for AI model providers.
A major CDN changed default crawler policy that affects publisher monetization, data access for AI model training, and crawler behavior industry-wide; it could materially alter how AI providers fetch and pay for web content.
Track Cloudflare Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Effective September 15, 2026, Cloudflare will by default block mixed-use crawlers from pages that host ads.
- The default change applies to new Cloudflare customers, new sites of existing customers, and all existing free customers unless site owners adjust settings.
- Cloudflare is evolving its Pay Per Crawl marketplace into a Pay Per Use model to let publishers charge AI companies when content generates value.
- Cloudflare is initially working with partners Ceramic.ai and You.com to pay publishers when content is used or appears in AI search results.
- Cloudflare data suggested over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages.
Connected Companies & Entities
4 Entities mapped“Cloudflare has just issued the AI industry a new deadline to separate the web crawlers used for traditional search purposes, like Google Sea...”
“Cloudflare specifically calls out the “world’s largest search engine” (clearly a Google reference!) as having access to about “2x more infor...”
“To put this into action, Cloudflare is initially working with two partners, Ceramic.ai and You.com....”
“Sarah has worked as a reporter for TechCrunch since August 2011....”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Cloudflare unveils AI-crawl controls, monetization tools
Cloudflare announced new classifications, analytics, and commercial partnerships to support an “agentic Internet” where automated agents and AI bots access web content under clearer rules. The company will test and finalize default settings over the next two months and plans to set defaults on September 15, 2026 that allow search but block training and agent use on pages with ads for new customers/sites (and for existing free customers who do not change settings). Cloudflare introduced an Attribution Business Insights dashboard, promoted Answer Engine Optimization (AEO) as a new discipline, and is evolving Pay Per Crawl into Pay Per Use to compensate publishers when their content creates value. It named partners including Ceramic.ai, You.com, beehiiv, Condé Nast and Patreon, and described prior initiatives such as AI Crawl Control and Web Bot Auth. Cloudflare says the changes aim to improve discoverability, reduce redundant crawling, and enable fairer monetization for creators.
Cloudflare introduces new controls for search crawlers, AI training, and agents.
Cloudflare is introducing new controls that let website operators separately decide whether their content can be used for traditional search, AI training, or AI agents. This aims to resolve the trade-off between being discoverable in search and protecting content from being used to train AI models. The company is replacing its 'Block AI Bots' toggle with three independent settings, including a 'Disallow AI Training' option. It is also introducing an 'Accountable' status for AI crawlers, which requires them to support independent opt-outs for AI training and search. Apple, Google, and Microsoft are already compliant or have committed to a timeline. Additionally, Cloudflare is replacing 'Managed Robots.txt' with 'Bot Preference Sync' for centralized control and is working on an open standard called 'ai-prefs' for expressing site preferences.
Cloudflare launches compliant crawler, sparking publisher tension
Cloudflare released a Crawl API (a crawl endpoint within its browser rendering API) that can scrape an entire website with one request and return content in HTML, Markdown, or structured JSON. The launch prompted publisher backlash after some sites reported they could not initially block Cloudflare’s crawler; Cloudflare product lead James Smith acknowledged messaging and implementation issues and said they have been fixed. The product is positioned as a compliant intermediary between publishers and AI builders, intended to respect publisher controls, reduce inefficient mass crawling, and create monetization options (following a prior pay-per-crawl offering). Publishers welcome tools that reduce server strain and preserve page performance, while some remain wary that intermediaries concentrating crawl control could shift power dynamics. Cloudflare says the goal is to establish best practices and support both supply (publishers) and demand (AI companies) sides of an emerging licensed AI content market.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
