Kunzum  /  Blog  /  Cloudflare AI Crawl Control: what it shows and what to allow

Cloudflare AI Crawl Control: what it shows and what to allow

Cloudflare AI Crawl Control shows which AI crawlers hit your site and sorts them by purpose, not just by category. That lets you allow search and live-fetch crawlers, treat training crawlers as a donation you can stop, and deny purposeless scrapers.

Agentic payments6 min read

Cloudflare AI Crawl Control: what it shows and what to allow

In short

  • Perplexity retrieved on 100% of edition-one answers (120 of 120), so blocking its live fetch removes every possible citation from that engine.
  • Claude retrieved on 65% (78 of 120) and ChatGPT on 54% (65 of 120), so blocking their live fetches drops you out of about two in three and one in two of their answers.
  • Cloudflare's June 2025 crawl-to-referral ratios were nearly 71,000 to 1 for Anthropic, 1,600 to 1 for OpenAI, 202.4 to 1 for Perplexity, 40 to 1 for Microsoft, and 9.4 to 1 for Google.
  • Charging crawlers is not a revenue line yet: TRM Labs screens only $25.62 million of x402's $52.7 million settled value as genuine commerce, so price training access if you sell that content anyway and leave the citing fetchers free.

This is part of the Kunzum reference on What is x402? The plain-English guide to agentic payments, which carries the figures this post draws on and the date each one was checked.

What AI Crawl Control actually shows

Cloudflare's AI Crawl Control, as described in Cloudflare's announcement, lets publishers see which AI crawlers reach a site and sort them by purpose, rather than blocking them wholesale. That sorting is the product. You get per-operator visibility and a purpose label on the agents hitting you, so a decision can be made per class of bot instead of per category.

Anything beyond that is not in the public documentation, and you should be suspicious of posts that describe dashboard widgets Cloudflare never shipped. What follows sticks to what the docs support.

The purposes, and which two can cite you

The documentation describes sorting by purpose rather than blocking wholesale. For a publisher, those purposes fall into three groups that behave very differently:

  • Model training. The bot collects text to train a model. Nothing it takes can produce a citation for your page, because there is no query being answered at that moment.
  • Search and index building. The bot builds an index that later answers queries. A page in that index can be retrieved and cited weeks later.
  • Live fetch on behalf of a user. The bot fetches your page while a person waits for an answer. This is the highest-intent crawl there is: the citation, if it comes, comes now.

The first trade is one-directional. You give text, you get nothing back in the answer layer. The second and third can put a page in front of a user, and only those two are worth defending.

What blocking each class costs

Kunzum measures whether answer engines retrieve at all on a fixed question set. The corpus holds 1,500 responses across two editions, and the retrieval rates below come from edition one, described in how the corpus was measured.

EngineRetrieved on edition-one answersAnswers a blocked fetcher could no longer cite you in
Perplexity100% (120 of 120)All of them
Claude65% (78 of 120)About two in three
ChatGPT54% (65 of 120)About one in two
Google AI OverviewsAlways retrievesEvery answer, which is why its rate is unset

Google AI Overviews always retrieves, so Kunzum does not report a rate for it. That cell is unset by design, not zero.

Read the table as a cost sheet. An engine can only cite a page it fetched, so the retrieval rate is the share of its answers a blocked fetcher drops you out of. Block Perplexity's user-triggered fetch and you leave the one engine in the set that retrieved on every measured answer. Block training crawlers and the table above does not move, because training crawlers do not appear in it at all.

Two honest limits. The corpus measures engines, not Cloudflare's purpose labels, so the step from "block this class" to "lose this much" is an inference, not a measurement. And each rate rests on 120 answers per engine. Small enough that a few points either way mean nothing.

A setup you can defend

Allow search and live-fetch crawlers. Allow training crawlers only if you have a reason to: model training is a donation, and you are allowed to stop donating. Deny the scrapers with no clear purpose behind them.

Then check it, because the number that decides this is the crawl-to-referral ratio, and Cloudflare publishes it. For the week of 19 to 26 June 2025, according to Cloudflare, crawl-to-referral ratios ran at nearly 71,000 to 1 for Anthropic, 1,600 to 1 for OpenAI, 202.4 to 1 for Perplexity, 40 to 1 for Microsoft and 9.4 to 1 for Google. Date those hard. They are 2025 figures, and the ratios have improved by more than an order of magnitude since.

Put a calendar note thirty days after you change anything. Same dashboard, same window, compare crawler mix and referral clicks. If referrals move and crawler volume does not, you changed nothing that matters. How to read x402 dashboards without being fooled applies here too: a large crawl count sitting next to a small referral number is a warning sign, not a win.

What charging crawlers is actually worth

Mostly nothing yet, and the numbers say why.

Cloudflare has two products pointed at this. Pay Per Crawl lets publishers price crawler access with Cloudflare as merchant of record, using HTTP 402, and it is in private beta according to Cloudflare. The Monetization Gateway, published 1 July 2026, goes wider: any protected asset, including pages, datasets, APIs and MCP tools, settled in stablecoins over x402.

The demand side is the problem. The x402 protocol charges no fee of its own; payers cover network fees. That makes the rail cheap, not the market large. According to TRM Labs, x402 has settled $52.7 million of total value across 198.9 million settlements since May 2025, and only $25.62 million screens as genuine commerce. Between 0.6% and 7.5% of that looks agentic. Citing Artemis, CoinDesk puts real daily x402 volume at around $28,000 against roughly 131,000 daily transactions and an average payment near $0.20.

Set that beside the volume of 402 responses Cloudflare says its network sends, bearing in mind that a 402 is any payment-required response, not an x402 offer. CoinDesk quotes Cloudflare's chief strategy officer at over a billion HTTP 402 responses a day. Cloudflare sits in front of 25.8% of all websites as of the September 2026 survey. A billion requests for payment a day in May 2026, against around $28,000 of daily x402 demand in March.

That gap is not obviously temporary. According to BlockEden.xyz, daily x402 transactions fell from about 731,000 in December 2025 to about 57,000 in February 2026, a decline of over 92%. Chainalysis reports well over 100 million cumulative x402 transactions on Base through Q1 2026, and payments of $1 or more rising from 49% of value in early 2025 to 95% by early 2026. That is consolidation into larger payments, which is the opposite of what a micropayment thesis needs.

So set up Pay Per Crawl if you want to be ready, and do not build a revenue line on it. Charge when a crawler is taking something you would otherwise sell, and when you can survive being dropped from the index that crawler feeds. Do not charge the live-fetch bots. Those are the ones that cite you.

If you want the wider map of what these engines do with the pages they retrieve, where AI engines cite crypto is the next thing to read.

Questions

What purposes does Cloudflare's AI Crawl Control sort AI crawlers by, and which ones can cite you?

AI Crawl Control sorts crawlers by purpose rather than blocking them wholesale. For a publisher those purposes fall into three groups: model training, search and index building, and live fetch on behalf of a user. Only the last two can put a page in front of a user and produce a citation.

What did Kunzum edition one find about retrieval rates for Perplexity, Claude, ChatGPT, and Google AI Overviews?

Perplexity retrieved on 100 percent of answers, 120 of 120. Claude retrieved on 65 percent, 78 of 120, and ChatGPT retrieved on 54 percent, 65 of 120. Google AI Overviews always retrieves, so Kunzum does not report a rate for it.

What crawl-to-referral ratios did Cloudflare report for the week of 19 to 26 June 2025?

Cloudflare reported nearly 71,000 to 1 for Anthropic, 1,600 to 1 for OpenAI, 202.4 to 1 for Perplexity, 40 to 1 for Microsoft, and 9.4 to 1 for Google. Those are 2025 figures, and the ratios have improved by more than an order of magnitude since.

Published 2026-10-01. Written by Narender Charan, who runs Kunzum.