Kunzum  /  Blog  /  Should you charge AI crawlers, block them, or leave them alone?

Should you charge AI crawlers, block them, or leave them alone?

You should block AI crawlers only when their output substitutes for your product or serving them costs more than it returns. Most sites should leave them alone or charge for access, because blocking an engine that rarely retrieves removes little while blocking a high-retrieval engine removes you from answers.

Agentic payments6 min read

Should you charge AI crawlers, block them, or leave them alone?

In short

  • Cloudflare's June 2025 crawl-to-refer ratios were nearly 71,000 to 1 for Anthropic, 1,600 to 1 for OpenAI, 202.4 to 1 for Perplexity, 40 to 1 for Microsoft and 9.4 to 1 for Google.
  • Kunzum's edition one retrieval rates were 100% for Perplexity (120 of 120), 65% for Claude (78 of 120) and 54% for ChatGPT (65 of 120); Google AI Overviews always retrieves but has no reported rate.
  • x402 settled roughly $52.7 million across 198.9 million settlements since May 2025, but real daily volume sits near $28,000 with an average payment near $0.20 across roughly 131,000 daily transactions.
  • Daily x402 transactions fell from about 731,000 in December 2025 to about 57,000 in February 2026, a decline of over 92%.

This is part of the Kunzum reference on What is x402? The plain-English guide to agentic payments, which carries the figures this post draws on and the date each one was checked.

The three options, stated plainly

Most publishers frame this as a moral question. It is a routing question: which automated readers get your pages, and what they pay for the privilege.

Leave them alone. This suits you if your revenue comes from being found rather than from licensing the finding. Documentation sites, price comparisons, tool pages, blogs that sell a service. The crawler costs bandwidth and returns a citation. Cloudflare shipped AI Crawl Control so publishers can see and sort AI crawlers by purpose instead of blocking them wholesale. That product only makes sense if most publishers should not block wholesale.

Charge them. This suits you if you hold something a model cannot synthesise from three other pages: a deep archive, a proprietary dataset, a real-time feed. Pay Per Crawl lets publishers price crawler access with Cloudflare as merchant of record, using HTTP 402. It was announced as a private beta. Cloudflare's Monetization Gateway, published 1 July 2026, extends the same mechanism to any protected asset, including pages, datasets, APIs and MCP tools, settling in stablecoins over x402.

Block them. This suits you if the crawler's output substitutes for your product, or if serving it is a real cost with no return. OpenAI's GPTBot and how to manage its web crawling behavior documents GPTBot and the robots.txt opt-out that sites use to block it. Paywalled newsrooms. Ticket marketplaces. Anything where the generated answer replaces the click.

If you cannot tell which of the three you are, you are probably the first.

What blocking actually costs

The argument for blocking is that crawlers take and never give back. Cloudflare has measured the take. In its analysis of crawl and refer ratios, for the week of 19 to 26 June 2025, it recorded ratios of nearly 71,000 to 1 for Anthropic, 1,600 to 1 for OpenAI, 202.4 to 1 for Perplexity, 40 to 1 for Microsoft and 9.4 to 1 for Google. Those are 2025 figures, and Cloudflare says the ratios have improved by more than an order of magnitude since. A blocking decision made on 2025 numbers is a decision made on stale numbers.

The same network sends over a billion HTTP 402 responses a day, according to Cloudflare's chief strategy officer, quoted by CoinDesk at Consensus in May 2026. That is a quoted remark rather than a published statistic, and it describes the network, not your site. Read it as evidence that the meter exists, not as evidence about your traffic.

Neither number tells you what blocking costs you. For that you need a different one: how often each engine retrieves at all.

Retrieval rates change the maths per engine

Blocking an engine that never searches the live web costs you nothing. Blocking one that searches on every answer removes you from every answer. The difference is not small.

Kunzum's own corpus measures this directly. It publishes 1,500 raw AI responses and the method behind them, including retrieval rates for edition one.

EngineRetrieval rate, edition oneWhat blocking its crawler removes
Perplexity100% (120 of 120)A retrieval on every answer
Claude65% (78 of 120)A retrieval on roughly two answers in three
ChatGPT54% (65 of 120)A retrieval on roughly half of answers
Google AI OverviewsAlways retrieves; rate not reportedA surface that cannot be measured the same way

Google's row is unset rather than zero. AI Overviews always retrieves, so there is no rate to report for it, and treating that blank as 0% is a misreading of the data.

Two caveats. Retrieval is not citation, and citation is not traffic. Kunzum measures whether the engine ran a web search, not whether your domain was the one it used. The wider citation picture, including which domains engines actually quote, sits in where AI engines cite crypto, drawn from 12,518 citations across 1,396 domains.

A blocking rule applied to an engine that never retrieves is a rule about nothing.

The paying side is thin, and that shapes the decision

Charging only works if someone buys. The x402 rail is real and the volume is small.

According to TRM Labs, roughly $52.7 million settled across 198.9 million x402 settlements on Base, Solana and Polygon since May 2025. That screens down to $25.62 million of plausible commerce, of which 0.6% to 7.5% appears agentic. According to CoinDesk, citing Artemis, real daily x402 volume sits at around $28,000, with an average payment near $0.20 across roughly 131,000 daily transactions. BlockEden.xyz reports daily x402 transactions falling from about 731,000 in December 2025 to about 57,000 in February 2026, a decline of over 92%.

So a per-crawl price is a bet on a demand curve that has not arrived at most sites. Cloudflare sits in front of 25.8% of all websites, as of the September 2026 W3Techs survey, which is why its defaults will shape the norm. Being able to charge is not the same as being paid.

How to measure your own position before deciding

Nobody can tell you from published data whether blocking pays at your site. That figure does not exist in public. You can produce it yourself in a week.

  1. Split your logs by user agent and count requests per crawler, not per engine. The categories do not line up.
  2. Work out what a human visit to the same pages is worth to you, using whatever conversion you already track. If you do not track it, that is the first problem.
  3. Block one engine, narrowly, for two weeks. Watch referral and citation counts separately, because they move on different clocks.
  4. Recheck the ratio data before extending the block. It moved by an order of magnitude in under a year.

If you sell something, start by leaving crawlers alone and reading what AI recommends instead to see whether you appear in the answers at all. If you do not appear, blocking is a decision about nothing. If you hold an archive nobody else has, price it and watch who pays.

For the mechanics, the free playbook covers the setup without the theory. The short version: leave crawlers alone unless you can name the revenue you are giving up, and measure that before you write the rule.

Questions

What are the three options for publishers handling AI crawlers?

Leave them alone, charge them, or block them. Leave them alone suits sites whose revenue comes from being found rather than licensing the finding, such as documentation sites, price comparisons, tool pages, and blogs that sell a service. Charging suits sites with a deep archive, a proprietary dataset, or a real-time feed. Blocking suits sites where the crawler's output substitutes for the product or where serving it is a real cost with no return.

What retrieval rates did Kunzum measure for Perplexity, Claude, and ChatGPT?

In Kunzum's edition one, Perplexity retrieved on 100% of answers, 120 of 120; Claude on 65%, 78 of 120; and ChatGPT on 54%, 65 of 120. Google AI Overviews always retrieves, but its rate was not reported. The post notes that treating that blank as 0% is a misreading of the data.

Why might a blocking decision based on Cloudflare's 2025 crawl and refer ratios be stale?

Cloudflare recorded those ratios for the week of 19 to 26 June 2025. They included nearly 71,000 to 1 for Anthropic and 1,600 to 1 for OpenAI. Cloudflare says the ratios have improved by more than an order of magnitude since, so a blocking decision made on 2025 numbers is a decision made on stale numbers.

Published 2026-09-20. Written by Narender Charan, who runs Kunzum.