Tool Teardowns
Teardown

I Checked 30 Sites for llms.txt. The AI Companies Don't Have One.

Aug 6, 2026 · 5 min read · by Jordan Kwan

TL;DR: For getting your site cited by AI assistants, llms.txt does nothing measurable: a 300,000-domain study found no correlation with citations, a 137,000-site log analysis found 97% of the files never receive a single request, and no AI vendor documents reading them. My own check makes the point harder: I fetched llms.txt from 30 prominent tech domains, and while 20 serve one, the main domains of OpenAI, Anthropic, Perplexity, xAI, and Google's AI properties all return 404. The one real use case is developer docs for coding agents. Everything else is decoration.

If you have touched SEO in the last year, someone has invoiced you, or tried to, for adding llms.txt. It is the perfect deliverable: cheap to generate, impossible for the client to evaluate, and blessed by a plausible story about "helping AI read your site." This site serves one, generated automatically at build time, so I have no anti-llms.txt axe to grind. I just wanted to know if the file does anything, and the evidence turns out to be unusually complete.

What is llms.txt supposed to do?

The proposal is real and reasonable. Jeremy Howard of Answer.AI proposed the standard in September 2024: a markdown file at /llms.txt giving language models a clean, structured index of your site, because "constructing the right context for LLMs based on a website is ambiguous" when everything is buried in HTML, navigation, and JavaScript. Note what the proposal actually targets: helping LLMs use a website at inference time, mostly documentation. It never promised search visibility. The promise inflation came later, from the people selling implementations.

Who actually serves one?

On August 6, 2026 I fetched https://domain/llms.txt from 30 prominent tech domains. 20 of 30 serve a real one: Stripe, Cloudflare, Shopify, GitHub, Vercel, Cursor, HubSpot, Slack, and most of the modern SaaS canon. And then the punchline: openai.com, anthropic.com, perplexity.ai, x.ai, and Google's AI domains all return 404. The five companies whose models would consume the file do not publish it where they live. (Fairness footnote: their developer docs subdomains mostly do serve one, which turns out to be the tell for what the file is actually for.) Also enjoyable: deepmind.google serves one while Google Search publicly says it ignores them, and langchain.com, a company that exists entirely to feed content to LLMs, does not bother.

Does anything actually read these files?

This is where the evidence gets brutal, and unlike the folklore stats this category usually runs on, these numbers have real sources. Ahrefs analyzed server logs across 137,210 domains in May 2026 and found that 97% of llms.txt files received zero requests. Of the 3% that got any traffic, the single largest requester category was SEO audit tools, at 21.7%, the industry checking its own homework, and AI retrieval bots accounted for 1.1%. Slackbot fetched llms.txt more often than PerplexityBot did. Ahrefs' authors put the conclusion in one sentence: "If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration."

The citation side agrees. SE Ranking's 300,000-domain study found no significant correlation between having the file and how often a domain gets cited in LLM answers; removing the llms.txt variable actually improved their model's accuracy. I also read the official crawler documentation for OpenAI, Anthropic, Perplexity, and Google: all four document robots.txt handling, and none mentions fetching llms.txt from third-party sites. Google's John Mueller called it comparable to the keywords meta tag: "AFAIK none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it)."

So why do 20 of 30 big tech companies serve one?

Three reasons, none of which is "it works for SEO." First, defaults: Mintlify made llms.txt automatic for every docs site it hosts in late 2024, instantly creating thousands of adopters. Second, cost: the file takes minutes to generate, so "why not" wins every internal debate, which is also why adoption grew 8.8x in a year while usage stayed at 3%. Third, and this is the legitimate one: developer docs for coding agents. The labs' own docs subdomains serve llms.txt because agents like Claude Code and Cursor genuinely do consume clean markdown docs at inference time, and in Ahrefs' data, AI agents were the healthiest slice of what little traffic the files get. If you publish API documentation, the file has a real audience. If you publish a blog, it has an audience of SEO audit tools.

Should you bother?

Keep it if it is free, skip it if it costs anything. This site's llms.txt is generated by the build pipeline at zero marginal effort, and that is the correct maximum spend. What actually moves AI citations, per every study in this post's sources, is the boring stuff: being crawlable at all, which is its own minefield, having extractable answer-shaped content, which most vendors fail on the one page that matters, since only 17 of 30 pricing pages state a price without JavaScript, and being the source of numbers other sites repeat. A proposal Jeremy Howard designed for inference-time docs became a GEO deliverable through pure incentive gravity: agencies needed something new to sell, and a file nobody reads is the easiest thing in the world to deliver. It is not the only such deliverable: across 26 GEO checklists, 70 of 90 recommended tactics were ordinary SEO and 17 had no evidence behind them at all. The 97% figure is not a maturity problem waiting to resolve. It is the market's answer.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗