Vivold Consulting
Business & Enterprise

Cloudflare to AI companies: separate search from scraping by September 15 - or get blocked

Default blocks on mixed-use crawlers plus a new Pay Per Use model hand content owners leverage - and hand AI builders a deadline

Key Insights

Cloudflare set a hard deadline: from September 15, 2026, its default settings will block mixed-use crawlers - those blending search, AI training, and agent traffic - from any pages carrying ads, unless site owners opt otherwise, with the change covering new customers, new sites, and all free-tier customers. CEO Matthew Prince pointedly called out the world's largest search engine for enjoying roughly 2x the information access of rivals by bundling search discoverability with AI harvesting, and revealed that bots now outnumber humans in internet traffic - a milestone that arrived a year early. Pay Per Crawl is evolving into Pay Per Use, paying publishers when content creates value in AI products, not merely when it's fetched.

Stay Updated

Get the latest insights delivered to your inbox

The web's traffic cop just changed the default

Cloudflare, which sits in front of a vast share of the world's websites, has issued the AI industry an ultimatum with a date attached. Starting September 15, 2026, its default settings will block mixed-use crawlers - bots that blend traditional search indexing with AI training and agent traffic - from any pages that host ads, unless the site owner deliberately opens the gate. The new defaults apply to new Cloudflare customers, new sites from existing customers, and all existing free customers, which in practice covers an enormous slice of the open web. The logic, per the company: most site owners want to be discoverable via search and even AI services, but they want protection against their intellectual property being harvested for free as the price of that discoverability.

The Google subtext, said out loud

Cloudflare's announcement barely disguises its target, noting that the world's largest search engine enjoys about twice the information access of other AI companies because remaining discoverable in its search index has been hard to separate from feeding its AI. Google has pushed back on that framing before, pointing to its Google-Extended control that lets sites opt out of training and AI products without losing Search placement - though its flagship Googlebot still crawls for Search features including AI Overviews and AI Mode, which is precisely the bundling Cloudflare wants unbundled. CEO Matthew Prince anchored the urgency in a startling milestone: the majority of internet traffic is now non-human, a threshold crossed roughly a year ahead of expectations.

From Pay Per Crawl to Pay Per Use

The stick comes with a marketplace. Cloudflare's existing Pay Per Crawl scheme - which lets sites charge AI bots for scraping - is evolving into Pay Per Use, under which publishers get paid when their content creates value inside an AI product, not merely when it is fetched. Initial partners Ceramic.ai and You.com will pay publishers when content surfaces in AI search results or when premium content is accessed, with the model open for other AI companies to adapt. There is an efficiency dividend too: Cloudflare's data shows over half of AI crawl traffic is wasted re-fetching unchanged pages, so structured commercial access could cut publishers' bandwidth bills even as it opens a revenue line.

Two to-do lists: content owners and AI builders

  • If you own content: audit your crawler traffic now, before the defaults flip. Decide deliberately - block, allow, or monetise - per bot category, and treat AI licensing as a nascent revenue line worth a pricing conversation. If your site runs on Cloudflare's free tier, the September change happens to you automatically; make it a choice instead.
  • If you build AI products or agents: your data-supply chain just acquired a deadline and a price tag. Inventory which sources your training refreshes and agent workflows depend on, budget for licensed access where it is critical, and make sure your crawlers are cleanly segmented by purpose - transparent, single-intent bots are exactly what this regime is designed to reward.
  • For marketers and strategists, the deeper shift is that visibility inside AI answers is becoming a negotiated, sometimes paid, channel rather than a free by-product of SEO. Build that assumption into 2027 content and distribution plans, and watch the Ceramic.ai and You.com pilots as the early template for how content gets compensated in an answer-engine world.

More in Business & Enterprise

All Business & Enterprise stories

Nadella's warning: you're paying for AI twice - once in tokens, once in your own IP

In a blog post, Satya Nadella warned that enterprises using proprietary AI models are paying twice - once in money for tokens, and again in the proprietary knowledge they must reveal to make those models useful, since models learn from the 'exhaust' of prompts, tool use, and especially corrections. He argued it is inconsistent for labs to claim fair-use rights to train on the world's public data while restricting others from distilling their models in return. Nadella - whose company invests in both OpenAI and Anthropic - later doubled down on CNN, saying firms without their own models or an AI gateway layer separating prompts, memory, and harness from the model won't survive as firms, having 'outsourced your thinking.'

Amazon retires Mechanical Turk: the platform that secretly powered 'AI' for 21 years is done

Amazon will close Mechanical Turk to new customers on July 30, 2026, moving the 21-year-old crowdsourcing marketplace into maintenance mode with no new features - and reporting indicates SageMaker Ground Truth and Amazon Augmented AI close to new customers the same day. Launched in 2005 as 'artificial artificial intelligence,' MTurk annotated the data that trained a generation of models; by 2023 a study found 33-46% of its workers were using LLMs to do the tasks, dissolving the platform's reason to exist. If your research, labeling, or human-review pipeline touches MTurk, you now have a migration deadline.

Zuckerberg's candid admission: AI agents 'haven't accelerated the way we expected'

At an internal town hall on July 2, Mark Zuckerberg told employees that AI agent development over the last four months has not accelerated as expected, that Meta's sweeping reorganisation was not as clean as it could have been, and that its bets on the new structure have not yet paid off - remarks first reported by Reuters from a recording. The admission stings because Meta laid off about 10% of its workforce and reassigned ~7,000 people to AI teams in May, with executives who planned the reorg reportedly optimistic about tools like Claude Code. Zuckerberg still expects significant AI benefits within three to six months - but the gap between agent hype and agent reality just got named by its biggest spender.