Is llms.txt worth it?
No, not as an AI visibility tactic. llms.txt is a proposed markdown file at your site root meant to give language models a curated map of your content. The idea is tidy, the file is harmless, and the measured evidence points one way: no major AI provider reads it for retrieval or training. This guide walks through that evidence in enough detail that you can check it yourself, because “trust us” is exactly the argument style this site exists to replace.
What was llms.txt supposed to do?
The proposal, published by Answer.AI in 2024, puts a clean markdown summary of your site at a predictable address: what the site is, which pages matter, where the good content lives. A language model could then skip your HTML noise, your navigation and cookie banners and scripts, and read the curated version. Reasonable in theory.
The proposal found real adoption among documentation sites and developer tools, and that visible adoption became the sales pitch: everyone’s adding it, so you should too. What the pitch skipped is the difference between sites serving the file and AI systems reading it.
What does the measured evidence show?
Two independent studies with different methods, one conclusion. Ahrefs analyzed logs across 137,000 domains and found that among the ~38,000 domains serving a valid llms.txt, 97% received zero requests for the file in a month. Not few requests. Zero. Otterly ran the controlled version: one site, correctly implemented file, 90 days of monitoring.
The site logged 62,100 AI bot visits in that window, of which 84 touched llms.txt, or about 0.1%, while an average content page drew roughly 265 bot visits. The file underperformed ordinary pages threefold. When the crawlers you’re trying to court visit your llms.txt less often than they visit a random blog post, the tactic has answered its own question.
Why don’t AI systems need the file?
Because they answer through live search. Ask ChatGPT a question and it rewrites your prompt into search queries, retrieves ranked results, and synthesizes from the pages it fetches. There is no step where the model detours to a site’s llms.txt to ask what it should read; the search index already answered that.
Pedro Dias, who spent six years on Google’s search quality team, explains the mismatch: llms.txt was designed for agentic browsing, meaning an AI-driven browser landing on a JavaScript-heavy app it cannot parse. At the retrieval layer it plays no role, and it contributes nothing to training because the file itself holds no content worth training on.
Why do tools keep scoring it anyway?
Because it’s easy to check, easy to sell, and sounds plausible in a sales call. A binary file-exists check costs a vendor nothing to build and gives customers a satisfying green checkmark to buy. That economic logic, rather than any retrieval logic, is why llms.txt persists in audit tools years after the log data came in.
It has become a genuinely useful tell in the other direction: when an agency’s proposal leads with llms.txt and schema markup as core deliverables, you have learned what you need to know about how they weigh evidence. Our scanner reports llms.txt as an unscored informational line, and the fix guide explains that choice to anyone who clicks.
Should you have one anyway?
Sure, if it costs you nothing, and for one narrow reason that is not the reason vendors sell. The honest position splits into two layers. For retrieval, meaning getting cited in AI answers, llms.txt does nothing, and no credible evidence says otherwise.
For agentic use, meaning an AI browsing agent landing on your site and needing a map of what lives where, the file can genuinely help, which is closer to what its creators designed it for in the first place. That use case is emerging as agent tooling spreads, it still earns you zero citations, and it is the reason we keep one ourselves, with that rationale written inside the file.
Adoption data makes the hype visible: in our scan of the web’s top 1,500 domains, 17.1% of top sites had an llms.txt, while the AI-SEO tool industry adopted it at 47.1%, nearly triple the rate, for a file whose measured retrieval effect is zero. We count ourselves in that 47.1% and score the file at zero anyway.
What you shouldn’t do is pay for it, prioritize it over anything on this list, or trust a score that awards points for it.
What actually moves AI visibility?
The boring, evidence-backed things. Confirm no firewall is blocking AI crawlers, since Cloudflare has blocked them by default on new domains since mid-2025 and that failure zeroes out everything else. Put your answer in the first 100 words, where Surfer measured almost 40% of AI citations originating.
Keep your content in the raw HTML, and rank for the questions people ask, because retrieval flows through search. Then run the free scan and see where you actually stand on the things that count.
See where your site stands. The free scan takes about fifteen seconds and shows every fix.
Check your real LLM visibility