How many websites block AI crawlers? We scanned the top 1,500 to find out
22.1% of the web’s biggest sites block at least one major AI crawler in robots.txt, and 9.5% block all five of them. That is what we found on August 25, 2026, when we ran AIOScan’s 21 documented checks against the 1,500 highest-ranked domains on the web. The blocking gets more aggressive the bigger the site: among the very top of the list, one in three sites blocks at least one AI crawler and more than one in five blocks all five. And access is not even the biggest problem. 45% of the sites we scanned scored a D or F on overall AI readiness, mostly because of how their content is shaped, not who they let in.
This page publishes the full numbers, the method, and the scores of the AI-SEO industry itself, ours included.
TL;DR: 22.1% of scanned top sites block at least one major AI crawler in robots.txt, 9.5% block all five, 14.2% block AI user agents at the firewall on top of that, llms.txt sits at 17.1% adoption (47.1% among AI-SEO vendors), and 45% of the web’s biggest sites grade D or F on AI readiness. Jump to a section:
- Method
- Robots.txt blocking rates per crawler
- The top-100 sites block twice as often
- Firewall-level blocking
- llms.txt adoption
- Content-shape failures and grades
- The AI-SEO vendor scoreboard
How did we run this study?
We scanned the top of the Tranco list, a research-grade ranking of the most popular domains on the web, on August 25, 2026.
We took the 1,500 highest-ranked domains after removing obvious infrastructure domains (CDNs, ad servers, DNS hosts), then ran each homepage through the same scanner you can use free on this site: one polite pass covering the page fetched with a normal user agent and again with an AI-crawler user agent, plus robots.txt, sitemap, and llms.txt. Every check and weight is published, so any number below can be independently re-derived.
Of the 1,500 domains, 893 served a scannable homepage. The rest failed DNS resolution from our test network, timed out, or refused automated requests outright, including roughly 200 that answered our self-identified scanner (AIOScanBot/1.0) with HTTP 403. Every percentage below is calculated over the 893 scanned sites, not the full 1,500, and the sample list is committed to our public repository so the study can be reproduced.
Independent work corroborates the access-layer numbers: HasData’s July 2026 index of 10,894 domains found closely comparable robots.txt blocking rates among top sites and documented the same gap between declared policy and firewall enforcement that we measure. Where this study differs is everything past the access layer: the answer-shape and readiness grades, the top-100 split, and the AI-SEO vendor scoreboard below exist nowhere else.
How many sites block AI crawlers in robots.txt?
22.1% of scanned sites block at least one of the five major AI crawlers, whether by naming it or through a blanket disallow rule that the crawler falls under. Per bot:
| AI crawler | Operator | Blocked by |
|---|---|---|
| CCBot | Common Crawl | 19.3% |
| GPTBot | OpenAI | 17.8% |
| ClaudeBot | Anthropic | 17.0% |
| Google-Extended | Google (Gemini training) | 16.0% |
| PerplexityBot | Perplexity | 13.1% |
9.5% of sites block all five at once. CCBot leads the block list, which matters more than most site owners realize: Common Crawl feeds the training data of nearly every major model. And a blanket User-agent: * / Disallow: / counts, because that is exactly how GPTBot reads it. Twitter’s wildcard disallow shuts GPTBot out without ever naming it, and LinkedIn goes further, banning all five crawlers by name. If you are blocking by accident, our guide to robots.txt for AI crawlers shows the exact rules to check.
Do the biggest sites block more?
Yes, at roughly double the rate. Among the very top of the ranking (Tranco rank 130 or better, 73 scannable sites), 32.9% block at least one AI crawler and 21.9% block all five, against 22.1% and 9.5% across the full sample. Their median AIO Score is also lower: 63 against 72 overall.
The pattern makes sense: the biggest platforms are walled gardens negotiating licensing deals, not publishers competing to be cited. For everyone else, that is the opportunity. The sites AI assistants can actually read and quote are disproportionately not the giants.
Do sites block AI crawlers beyond robots.txt?
14.2% of scanned sites refused or challenged a request carrying an AI-crawler user agent even though we could fetch the same page as a normal visitor. This is the firewall layer: bot-protection rules, often default CDN or WAF settings, that no robots.txt audit will ever surface. It is also the most invisible failure in AI visibility, because everything looks fine in a browser while GPTBot, ClaudeBot, and PerplexityBot get turned away at the door. We check it by actually fetching your page with an AI-crawler user agent, and our firewall fix guide covers the settings to change in Cloudflare and the other major CDNs.
How many websites have an llms.txt file?
17.1% of the top sites we scanned have an llms.txt file. The AI-SEO tool industry itself adopts it at 47.1%, nearly three times the rate, which is a remarkable number for a file that Ahrefs’ 137,000-domain log study found receives zero requests on 97% of the domains that host it. The industry is following its own hype rather than its own evidence. Full disclosure: we are part of that 47.1%. aioscan.com keeps an llms.txt for agent navigation, the one use case with a real mechanism behind it, states that rationale inside the file, and scores the file at zero for visibility, ours included.
What actually fails: the content, not the access
Access problems are real but minority problems. Content shape fails on most of the web’s biggest sites:
- 71.8% have no question-formatted headings anywhere on their homepage
- 44.8% fail the answer-up-front test outright, burying their first direct answer below the first 100 words
- 53.6% have no valid structured data on the homepage
- The overall grade distribution: 1.1% A+, 8.7% A, 23.2% B, 21.9% C, 15.2% D, 29.8% F
The median top-1,500 site scores 72 of 100. These are the pages with the most authority on the web, and most of them are still written for a reader who will scroll, not for a model that quotes the first clean answer it finds. That gap is exactly what smaller sites can exploit, because AI crawlers read pages differently than human visitors do, and shaping content for them costs nothing but editing. Question headings are the highest-value fix on that list, and the fix library covers it step by step.
How does the AI-SEO industry itself score?
We also scanned 20 companies that sell AI visibility, SEO, or GEO tooling, including us. Full disclosure before the table: AIOScan scores 100 on its own scanner, which should surprise nobody, since we built the site to pass the test we publish.
The honest value of this table is everyone else’s row, scanned with the same public methodology on the same day, and any vendor can rerun the free scan to verify or dispute their number.
| Site | AIO Score | Grade | Has llms.txt | Answer up front |
|---|---|---|---|---|
| aioscan.com (ours) | 100 | A+ | Yes | Pass |
| conductor.com | 95 | A | Yes | Pass |
| hubspot.com | 95 | A | Yes | Pass |
| semrush.com | 93 | A | Yes | Pass |
| surferseo.com | 92 | A | No | Pass |
| searchatlas.com | 88 | B | Yes | Partial |
| ahrefs.com | 85 | B | No | Partial |
| writesonic.com | 85 | B | Yes | Pass |
| llmrefs.com | 84 | B | Yes | Partial |
| seranking.com | 84 | B | No | Partial |
| moz.com | 83 | B | No | Partial |
| screamingfrog.co.uk | 83 | B | No | Pass |
| sistrix.com | 78 | C | No | Pass |
| mangools.com | 77 | C | No | Pass |
| brightedge.com | 75 | C | Yes | Fail |
| peec.ai | 72 | C | No | Fail |
| tryprofound.com | 72 | C | No | Partial |
Not a single scanned vendor blocks any AI crawler in robots.txt, which tells you where the industry’s incentives point. Three vendors could not be scored: otterly.ai answered our self-identified scanner with HTTP 403 on repeated attempts, and similarweb.com and scrunch.ai did not resolve from our test network, so we make no claims about any of them. The industry averages a respectable 85, but a third of the vendors selling AI visibility sit in the C range on a published, reproducible test of it.
What should site owners take from this?
Check access first, because it is binary: one bad firewall rule or robots line makes every other optimization irrelevant. Then fix content shape, because that is where most of the web is leaving citations on the table and where the competition is thinnest.
Both take minutes to diagnose: the same 21 checks this study ran are free to run on your own site, with a fix guide attached to anything that fails.
We will rerun this study and update the numbers as the landscape shifts. If you use these figures, cite this page: the sample, method, and aggregates are public, and the numbers are only as good as their reproducibility.
See where your site stands. The free scan takes about fifteen seconds and shows every fix.
Run a free AI visibility scan