Ranking on Google and Invisible to ChatGPT: I Read 24 robots.txt Files to Find Out Why
Aug 6, 2026 · 5 min read · by Jordan Kwan
TL;DR: Google visibility and ChatGPT visibility are different systems with different gates, and passing one says nothing about the other. ChatGPT search leans on Bing's index plus OpenAI's own search crawler, engines barely overlap in what they cite (11% of domains shared between ChatGPT and Perplexity), and a thick layer of the web has simply locked AI crawlers out, deliberately or by inherited default. I read the robots.txt of 24 major sites to map who blocks what: 8 block OpenAI's training bot, only 5 block its search bot, and the sites that know what they are doing block one without the other.
A site owner's 2026 complaint, everywhere from r/seogrowth to agency sales decks: we rank fine on Google, and ChatGPT acts like we do not exist. This is not a mystery or an algorithm grudge. It is three mechanical differences, each checkable, and I checked them.
How can you rank on Google and not exist in ChatGPT?
Different index. When ChatGPT searches the web, OpenAI documents Bing as a search provider alongside its own crawler, and the correlation is dramatic: Seer Interactive matched 87% of SearchGPT citations to Bing's top organic results for the same queries, versus a 56% match against Google. Your Google rank is invisible to that pipeline; your Bing indexation is the pipeline. A site that never bothered with Bing Webmaster Tools, which describes most sites, is running this race with one shoe.
Different engines, different sources. Profound ran 100,000 identical prompts through ChatGPT and Perplexity and found only 11.0% of cited domains appeared in both; in the authors' words, "nearly 89% of AI citations come from completely different sources depending on which model users query." Being cited by one assistant is not a rising tide. Each engine is its own distribution channel with its own habits.
Different gate. And then there is the blunt one: the crawler cannot cite what it is forbidden to read.
Who blocks what? The robots.txt census
On August 6, 2026 I fetched robots.txt from 24 major domains: national news sites, tech media, review platforms, community sites, and B2B SaaS. The tallies: 8 of 24 block GPTBot (OpenAI's training crawler), 5 block OAI-SearchBot (its search crawler), 10 block PerplexityBot, 12 block ClaudeBot, and 3 more (Reddit, Stack Overflow, Reuters) block everything by wildcard. That is consistent with the bigger trackers: Ben Welsh's news-publisher census has 48.1% of 1,153 outlets blocking OpenAI's crawler as of this month, and Otterly's vendor analysis of a million citations claims 73% of sites carry some technical barrier to AI crawler access.
The pattern inside the pattern is the finding. Four sites block the training bot while leaving the search bot open, and zero do the reverse. Reuters is the cleanest: it blocks everything by default, then explicitly allowlists OAI-SearchBot and ChatGPT-User, a surgical "no training on our archive, yes to being cited in answers" posture. Meanwhile the licensing economy is legible in the files themselves: WSJ and The Verge, both under OpenAI deals, explicitly allow GPTBot while The Verge still blocks Perplexity, Claude, and Google-Extended. Robots.txt has quietly become where AI business strategy is written down, and where it stops: 22 of 27 paywalled publishers block Anthropic's training crawler and only 13 block the agent that fetches on a reader's command.
The other side of the census: the B2B half (HubSpot, Salesforce, Shopify, Stripe, Wikipedia) names no AI bots at all, fully open. If you sell software, your competitors' content is almost certainly crawlable, which makes an accidental block on your side a pure gift to them.
Why "accidental" is the operative word
Since July 2025, Cloudflare blocks AI crawlers by default for every new domain, with CEO Matthew Prince framing it as giving "publishers the control they deserve." Whatever you think of the economics (by July 2026 the model had already evolved into paying publishers per answer rather than per crawl), the practical effect is that thousands of sites inherited an AI block as a CDN default, not a decision. A site can be Google-visible, Bing-indexed, and still invisible to assistants because of a firewall toggle nobody on the team has ever seen. Rand Fishkin's line about AI answers applies with no appeal process: "If you're invisible, you're invisible."
What should you check on your own site?
Four checks, in causal order. One: fetch your own robots.txt and your CDN's bot settings, looking for GPTBot, OAI-SearchBot, PerplexityBot, and ClaudeBot by name, and decide each one on purpose; blocking training while allowing search is a coherent position, but only if you chose it. Two: verify Bing indexation in Bing Webmaster Tools, the unglamorous prerequisite (this site's own setup started there) that the 87% correlation makes non-optional. Three: confirm your pages are actually extractable, since a decorative llms.txt does not substitute for crawlable HTML. Four: instrument the outcome in your analytics' AI channel, because per-engine visibility at 11% overlap means you have to measure each channel to know any of them, and because the shortcut metric everyone quotes instead is unstable: 26 articles attribute 22 different crawl-to-refer ratios to Anthropic alone.
What does this not prove?
The census is 24 hand-picked major domains, not a random sample; the true blocking rates for the broader web sit somewhere between my tallies and Otterly's 73%, and that spread is wide. The 87% Bing correlation is from early 2025 and OpenAI's own index has grown since, so treat Bing as the dominant gate, not the only one. What the record does establish is the answer to the headline question: nothing about ranking on Google earns you anything in an AI answer, because the index, the citation habits, and the gate are all different, and at least one of the three is usually closed without the owner knowing.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗