Is llms.txt worth it?
No, not as an AI visibility tactic. llms.txt is a proposed markdown file at your site root meant to give language models a curated map of your content. The idea is tidy, the file is harmless, and the measured evidence points one way: no major AI provider reads it for retrieval or training. This guide walks through that evidence in enough detail that you can check it yourself, because “trust us” is exactly the argument style this site exists to replace.
What was llms.txt supposed to do?
The proposal, published by Answer.AI in 2024, puts a clean markdown summary of your site at a predictable address: what the site is, which pages matter, where the good content lives. A language model could then skip your HTML noise, your navigation and cookie banners and scripts, and read the curated version. Reasonable in theory.
The proposal found real adoption among documentation sites and developer tools, and that visible adoption became the sales pitch: everyone’s adding it, so you should too. What the pitch skipped is the difference between sites serving the file and AI systems reading it.
What does the measured evidence show?
Two independent studies with different methods, one conclusion. Ahrefs analyzed logs across 137,000 domains and found that among the ~38,000 domains serving a valid llms.txt, 97% received zero requests for the file in a month. Not few requests. Zero. Otterly ran the controlled version: one site, correctly implemented file, 90 days of monitoring.
The site logged 62,100 AI bot visits in that window, of which 84 touched llms.txt, or about 0.1%, while an average content page drew roughly 265 bot visits. The file underperformed ordinary pages threefold. When the crawlers you’re trying to court visit your llms.txt less often than they visit a random blog post, the tactic has answered its own question.
Why don’t AI systems need the file?
Because they answer through live search. Ask ChatGPT a question and it rewrites your prompt into search queries, retrieves ranked results, and synthesizes from the pages it fetches. There is no step where the model detours to a site’s llms.txt to ask what it should read; the search index already answered that.
Pedro Dias, who spent six years on Google’s search quality team, explains the mismatch: llms.txt was designed for agentic browsing, meaning an AI-driven browser landing on a JavaScript-heavy app it cannot parse. At the retrieval layer it plays no role, and it contributes nothing to training because the file itself holds no content worth training on.
Why do tools keep scoring it anyway?
Because it’s easy to check, easy to sell, and sounds plausible in a sales call. A binary file-exists check costs a vendor nothing to build and gives customers a satisfying green checkmark to buy. That economic logic, rather than any retrieval logic, is why llms.txt persists in audit tools years after the log data came in.
It has become a genuinely useful tell in the other direction: when an agency’s proposal leads with llms.txt and schema markup as core deliverables, you have learned what you need to know about how they weigh evidence. Our scanner reports llms.txt as an unscored informational line, and the fix guide explains that choice to anyone who clicks.
Should you have one anyway?
Sure, if it costs you nothing. We keep one ourselves. It’s harmless, some documentation tools consume the format for their own purposes, and if adoption ever genuinely changes we would rather be a file ahead than behind. We’d also change our scoring in the open, with the new evidence linked, exactly as we did when the current evidence arrived.
What you shouldn’t do is pay for it, prioritize it over anything on this list, or trust a score that awards points for it.
What actually moves AI visibility?
The boring, evidence-backed things. Confirm no firewall is blocking AI crawlers, since Cloudflare has blocked them by default on new domains since mid-2025 and that failure zeroes out everything else. Put your answer in the first 100 words, where Surfer measured almost 40% of AI citations originating.
Keep your content in the raw HTML, and rank for the questions people ask, because retrieval flows through search. Then run the free scan and see where you actually stand on the things that count.
See where your site stands. The free scan takes about fifteen seconds and shows every fix.
Check your real LLM visibility