Research
Research

I Asked 40 Sites for Their Homepage as GPTBot. Seven Said 402 Payment Required. None Said How Much.

Aug 15, 2026 · 8 min read · by Jordan Kwan

TL;DR: On August 15, 2026 I requested 40 homepages four times each, once as a browser and once each as GPTBot, ClaudeBot and Googlebot, and logged the status and response size of all 160 requests. Seven of the 40 returned 402 Payment Required to at least one AI user-agent: forbes.com, rollingstone.com, variety.com, usatoday.com, people.com, theatlantic.com and independent.co.uk. Three of those seven charged one AI operator and not the other. Not one of the seven sent the crawler-price header that Cloudflare's pay-per-crawl spec defines, and four of the seven came from TollBit, not Cloudflare. In the wild, a 402 today is a locked door with no price written on it.

Which sites return 402 to an AI crawler?

Cloudflare's pay per crawl revived HTTP 402 as a billing primitive in July 2025, and on September 15, 2026 its defaults change to block training and agent crawlers on ad-carrying pages for new domains. Everyone covered the announcement. Nobody published what the web returns today.

So I asked it. 40 domains: 25 publishers, 10 large e-commerce and review sites, 5 AI companies. Each got four GET requests to its root URL, redirects followed, body discarded, via curl -s -L -o /dev/null -w "%{http_code} %{size_download}". Sizes are compressed transfer bytes. 151 of 160 requests returned an HTTP status. Nine did not: washingtonpost.com, bestbuy.com and ebay.com each timed out or dropped the connection for three of their four user-agents, identically on a retry, so those cells stay empty rather than get quietly refilled.

Domain Browser UA GPTBot ClaudeBot Googlebot
forbes.com 200 (134k) 402 (156B) 402 (156B) 200 (135k)
rollingstone.com 200 (124k) 402 (156B) 402 (156B) 200 (124k)
people.com 200 (92k) 403 (328k) 402 (109B) 460 (10B)
theatlantic.com 200 (79k) 200 (79k) 402 (55B) 200 (79k)
independent.co.uk 200 (56k) 200 (56k) 402 (4k) 200 (56k)
nytimes.com 200 (265k) 403 (139k) 403 (139k) 200 (201k)
wsj.com 401 (770B) 401 (767B) 401 (767B) 401 (767B)
washingtonpost.com no response no response 403 (376B) no response
theguardian.com 200 (134k) 200 (134k) 403 (20k) 200 (134k)
reuters.com 401 (773B) 401 (771B) 401 (771B) 401 (771B)
bloomberg.com 403 (13k) 403 (13k) 403 (13k) 403 (13k)
cnn.com 200 (582k) 451 (38B) 451 (38B) 403 (424B)
bbc.com 200 (92k) 200 (92k) 200 (92k) 200 (92k)
time.com 200 (155k) 200 (10k) 200 (10k) 200 (155k)
wired.com 200 (180k) 200 (180k) 200 (180k) 200 (180k)
theverge.com 200 (63k) 200 (63k) 403 (11B) 200 (63k)
techcrunch.com 200 (57k) 200 (57k) 200 (57k) 200 (57k)
arstechnica.com 200 (41k) 403 (118B) 200 (41k) 200 (41k)
businessinsider.com 200 (60k) 200 (60k) 200 (60k) 200 (60k)
vox.com 200 (39k) 200 (39k) 403 (11B) 200 (39k)
newyorker.com 200 (192k) 200 (192k) 403 (919B) 200 (192k)
usatoday.com 200 (76k) 402 (235B) 402 (235B) 200 (76k)
nypost.com 200 (159k) 200 (159k) 200 (159k) 200 (159k)
variety.com 200 (98k) 402 (156B) 402 (156B) 200 (98k)
espn.com 202 (0B) 200 (45k) 200 (44k) 202 (0B)
amazon.com 200 (1k) 200 (1k) 200 (1k) 200 (1k)
walmart.com 200 (67k) 200 (67k) 200 (67k) 200 (67k)
target.com 200 (81k) 200 (60k) 200 (60k) 200 (61k)
bestbuy.com no response no response no response 403 (369B)
etsy.com 403 (776B) 403 (776B) 403 (776B) 429 (337B)
ebay.com 200 (87k) no response no response no response
wayfair.com 429 (5k) 429 (5k) 429 (5k) 429 (5k)
yelp.com 403 (779B) 403 (0B) 403 (776B) 403 (776B)
tripadvisor.com 403 (517B) 403 (516B) 403 (506B) 403 (516B)
g2.com 403 (909B) 403 (913B) 403 (911B) 403 (911B)
openai.com 200 (157k) 403 (5k) 403 (5k) 403 (5k)
anthropic.com 200 (41k) 200 (41k) 200 (41k) 200 (41k)
perplexity.ai 200 (3k) 403 (6k) 403 (6k) 403 (3k)
cohere.com 200 (73k) 200 (73k) 200 (73k) 200 (73k)
mistral.ai 200 (81k) 200 (81k) 200 (81k) 200 (81k)

Fifteen domains 403'd an AI user-agent, but five of those (bloomberg.com, etsy.com, yelp.com, tripadvisor.com, g2.com) 403'd my plain browser string too, so that is generic bot mitigation, not AI policy. Every one of these is a decision somebody made by hand, which is the opposite of what happens on hosted platforms, where 24 of 29 name no AI crawler at all. Ten refusals were AI-specific. Only 29 of 40 served a 200 to the browser string at all.

What is actually sending the 402?

I re-requested all seven and read the headers, because a status code is not evidence of a business model. Three different stacks produce that number.

Four are TollBit. forbes.com, rollingstone.com and variety.com 307-redirect AI user-agents to a tollbit. subdomain, where a Caddy server returns 402 with X-Tollbit-Forwarded: true and the body You are not authorized to access this content without a valid TollBit Token. usatoday.com returns that payload inline. TollBit is a separate vendor from Cloudflare selling publishers a bot paywall, and it, not pay-per-crawl, is most of what the 402 web currently is. The standards-track alternative is doing worse: of 36 organizations named as backers of the RSL licensing standard, three actually serve the license file.

Two are Cloudflare's edge, and still not pay-per-crawl. people.com and theatlantic.com answer with Server: cloudflare and a CF-RAY. The Atlantic's body is {"message":"Please contact the site owner for access."}; People's is a licensing email address. Cloudflare's docs say pay per crawl is in closed beta and that an unpaid crawler gets a 402 "with pricing," carried in a crawler-price header. Zero of the seven sent crawler-price. These are custom rules borrowing the status code, not transactions.

One is the publisher's own origin. independent.co.uk returns a branded page titled "Payment Required | The Independent" through Varnish, with X-Backend: flow-uk and the nonstandard reason phrase 402 Unknown Error.

Does the answer depend on which AI company is asking?

On three of 40, yes. people.com returns 403 to GPTBot and 402 to ClaudeBot. theatlantic.com and independent.co.uk serve GPTBot a full 200 and bill ClaudeBot. Reverse it and arstechnica.com 403s GPTBot while handing ClaudeBot the identical 41k page. Counting any status divergence between the two AI strings, 8 of 40 treat OpenAI and Anthropic differently, setting aside washingtonpost.com where three of four requests never completed. The robots.txt layer already varied by operator; this is the same discrimination enforced in the response, live. I re-ran the nine divergent domains with the exact GPTBot/1.4 string OpenAI currently documents next to the 1.2 string I had used, and every status matched, so these rules key on the substring, not the version.

Is anybody actually charging, then?

One site, and it never uses 402. time.com returns 200 to everything, but AI strings get 10k where a browser gets 155k: a 93% smaller payload, and the only material Googlebot-versus-GPTBot divergence in 18 comparable pairs. That response is content-type: text/markdown, opens with <!-- mobian-agent-page publisher="time" -->, and carries x-mobian-tokens: 10481 plus an x-mobian-impression UUID. TIME is not gating the crawler. It is metering it and selling ad inventory inside the markdown via Mobian, whose agent ads TIME launched in July 2026 with Ally Bank as a first buyer. The 402 crowd turns crawlers away. TIME counted their tokens and billed someone else. If you have treated a markdown file for agents as a formality, that token header is the counterexample.

What does this method not prove?

Three things, and they matter more than the table.

Sending a user-agent string is not being that crawler. OpenAI publishes 21 IPv4 prefixes in openai.com/gptbot.json; Anthropic publishes a list at claude.com/crawling/bots.json and says a source IP on it is what identifies its crawler. My requests came from a residential Telus IP in North Vancouver, on neither list. A site verifying by IP would treat real GPTBot traffic differently. So read every cell as what the edge returns to anything claiming to be GPTBot, which is the interesting part: these edges are pricing an unverified string, and an unverified string is free to type.

A 402 or 403 is not proof of a pay-per-crawl configuration. I could separate TollBit, Cloudflare and origin rules only because the headers said so. Cloudflare's own content signals policy, pushed to over 3.8 million domains, concedes that signals "are not technical countermeasures against scraping." A status code is posture until a price rides along with it.

One vantage point, one clock. Every number here is one residential connection in Canada between 05:24 and 05:32 UTC on August 16, 2026. Geo-routing, cache state and staged rollouts move these cells, and the September 15 default flip will move them again. Re-run it from your own network before quoting it, and instrument your own AI traffic rather than inferring it from mine.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗