An Index — kept by Alek

Dataset & Analysis

What the web's busiest sites refuse to let you crawl

We asked the top 1,000,000 domains on the internet for their robots.txt, parsed every rule, and counted what they allow, what they block, and — increasingly — who they block.

1,000,000
domains queried
Tranco top 1,000,000
60.5%
served a real robots.txt
605,046 parseable files
17.0%
block at least one AI crawler
103,045 of 605,046
1.7×
more sites name GPTBot than Googlebot
119,287 vs 69,422

01 — Coverage

Asking 1,000,000 sites a simple question

A request for /robots.txt is the most basic thing a crawler does. Only 60.5% of the top 1,000,000 domains answered it with an actual robots file.

The rest split into instructive failure modes. Some domains in a popularity ranking aren't websites at all — CDN edges, nameserver infrastructure, and tracking hosts inflate the DNS failures. Others actively refuse: 74,194 returned an HTTP error, most of them 403 from bot-protection layers that block the very file designed to tell bots what to do. And 47,958 served an HTML page with a 200 status — a soft 404 that a naive parser would happily read as a robots file with zero rules.

Served robots.txt: 605,046 (60.5%) 60.5% Network / DNS failure: 164,682 (16.5%) 16.5% HTTP error (403, 5xx…): 74,194 (7.4%) 7.4% 404 / 410 — no file: 100,591 (10.1%) 10.1% Unfollowable redirect: 69 (0.0%) HTML page, not robots: 47,958 (4.8%) Empty file: 7,460 (0.7%)
  • Served robots.txt 605,046 60.5%
  • Network / DNS failure 164,682 16.5%
  • HTTP error (403, 5xx…) 74,194 7.4%
  • 404 / 410 — no file 100,591 10.1%
  • Unfollowable redirect 69 0.0%
  • HTML page, not robots 47,958 4.8%
  • Empty file 7,460 0.7%
Outcome of one GET request to /robots.txt per domain, HTTPS with an HTTP fallback and a www. retry for domains whose apex does not resolve.
The parsing trap. 4.8% of all domains return HTML dressed as text/plain-shaped content. Treat those as valid and you conclude they permit everything — the exact opposite of what many of them intend.

02 — Default posture

Most of the web is open, conditionally

For the generic User-agent: * group — the rules that apply to a crawler nobody has heard of — the dominant answer is yes, but not there.

Open — no restrictions: 173,246 (28.6%) 28.6% Partial — some paths blocked: 344,390 (56.9%) 56.9% Blocked — Disallow: /: 48,681 (8.0%) 8.0% No wildcard group: 38,729 (6.4%)
  • Open — no restrictions 173,246 28.6%
  • Partial — some paths blocked 344,390 56.9%
  • Blocked — Disallow: / 48,681 8.0%
  • No wildcard group 38,729 6.4%
Stance of the wildcard group among the 605,046 parseable robots.txt files. “Blocked” means Disallow: / with no Allow carve-out.

Only 48,681 sites (8.0%) slam the door on unknown crawlers entirely. The typical file is a scalpel, not a wall: a median of 8 rules across 1 user-agent group, carving out the parts of a site that are expensive to serve or worthless to index.

03 — The AI question

The new names on the blocklist

The most significant change to robots.txt in twenty years is not a new directive. It's a new class of user-agent — and site owners are naming them explicitly.

154,160 sites (25.5% of those serving a robots file) name at least one AI or LLM-training crawler by name. Of those, the overwhelming majority are not writing rules to welcome them: 103,045 sites block at least one outright.

names the bot blocks it outright CCBot — named by 119,287 sites, blocked by 79,830 (66.9% of those that name it) CCBot Common Crawl 66.9% of 119,287 Bytespider — named by 112,269 sites, blocked by 76,212 (67.9% of those that name it) Bytespider ByteDance 67.9% of 112,269 GPTBot — named by 102,412 sites, blocked by 80,957 (79.1% of those that name it) GPTBot OpenAI 79.1% of 102,412 ClaudeBot — named by 93,580 sites, blocked by 74,561 (79.7% of those that name it) ClaudeBot Anthropic 79.7% of 93,580 Google-Extended — named by 86,346 sites, blocked by 69,755 (80.8% of those that name it) Google-Extended Google 80.8% of 86,346 Amazonbot — named by 84,408 sites, blocked by 74,037 (87.7% of those that name it) Amazonbot Amazon 87.7% of 84,408 Meta-ExternalAgent — named by 79,302 sites, blocked by 67,078 (84.6% of those that name it) Meta-ExternalAgent Meta 84.6% of 79,302 Applebot-Extended — named by 76,643 sites, blocked by 67,021 (87.4% of those that name it) Applebot-Extended Apple 87.4% of 76,643 PetalBot — named by 49,836 sites, blocked by 17,474 (35.1% of those that name it) PetalBot Huawei 35.1% of 49,836 ChatGPT-User — named by 32,649 sites, blocked by 15,967 (48.9% of those that name it) ChatGPT-User OpenAI 48.9% of 32,649 PerplexityBot — named by 31,283 sites, blocked by 13,070 (41.8% of those that name it) PerplexityBot Perplexity 41.8% of 31,283 anthropic-ai — named by 26,444 sites, blocked by 16,114 (60.9% of those that name it) anthropic-ai Anthropic 60.9% of 26,444 OAI-SearchBot — named by 21,643 sites, blocked by 6,409 (29.6% of those that name it) OAI-SearchBot OpenAI 29.6% of 21,643 Claude-Web — named by 19,844 sites, blocked by 12,607 (63.5% of those that name it) Claude-Web Anthropic 63.5% of 19,844 cohere-ai — named by 19,579 sites, blocked by 12,756 (65.2% of those that name it) cohere-ai Cohere 65.2% of 19,579 FacebookBot — named by 17,050 sites, blocked by 10,187 (59.7% of those that name it) FacebookBot Meta 59.7% of 17,050 YouBot — named by 16,709 sites, blocked by 10,570 (63.3% of those that name it) YouBot You.com 63.3% of 16,709 omgili — named by 16,372 sites, blocked by 12,666 (77.4% of those that name it) omgili Webz.io 77.4% of 16,372

When a site bothers to mention an AI crawler at all, the mention is usually a refusal — a block rate far above anything the traditional search crawlers see.

CrawlerOperatorSites naming it% of corpusBlocked byBlock rate
CCBotCommon Crawl119,28719.72%79,83066.9%
BytespiderByteDance112,26918.56%76,21267.9%
GPTBotOpenAI102,41216.93%80,95779.1%
ClaudeBotAnthropic93,58015.47%74,56179.7%
Google-ExtendedGoogle86,34614.27%69,75580.8%
AmazonbotAmazon84,40813.95%74,03787.7%
Meta-ExternalAgentMeta79,30213.11%67,07884.6%
Applebot-ExtendedApple76,64312.67%67,02187.4%
PetalBotHuawei49,8368.24%17,47435.1%
ChatGPT-UserOpenAI32,6495.4%15,96748.9%
PerplexityBotPerplexity31,2835.17%13,07041.8%
anthropic-aiAnthropic26,4444.37%16,11460.9%
OAI-SearchBotOpenAI21,6433.58%6,40929.6%
Claude-WebAnthropic19,8443.28%12,60763.5%
cohere-aiCohere19,5793.24%12,75665.2%
FacebookBotMeta17,0502.82%10,18759.7%
YouBotYou.com16,7092.76%10,57063.3%
omgiliWebz.io16,3722.71%12,66677.4%
omgilibotWebz.io16,0482.65%12,22876.2%
ImagesiftBotImagesift15,1472.5%12,24180.8%
DiffbotDiffbot14,2392.35%10,45173.4%
TimpibotTimpi10,1891.68%7,55874.2%

04 — For contrast

Nobody blocks Googlebot

Run the same measurement against the crawlers that have been around for decades and the picture inverts.

names the bot blocks it outright AhrefsBot — named by 69,836 sites, blocked by 25,044 (35.9% of those that name it) AhrefsBot seo 35.9% of 69,836 Googlebot — named by 69,422 sites, blocked by 814 (1.2% of those that name it) Googlebot search 1.2% of 69,422 MJ12bot — named by 67,310 sites, blocked by 28,565 (42.4% of those that name it) MJ12bot seo 42.4% of 67,310 SemrushBot — named by 59,584 sites, blocked by 24,934 (41.8% of those that name it) SemrushBot seo 41.8% of 59,584 Bingbot — named by 56,092 sites, blocked by 1,693 (3.0% of those that name it) Bingbot search 3.0% of 56,092 DotBot — named by 53,558 sites, blocked by 21,198 (39.6% of those that name it) DotBot seo 39.6% of 53,558 YandexBot — named by 37,626 sites, blocked by 5,687 (15.1% of those that name it) YandexBot search 15.1% of 37,626 rogerbot — named by 35,595 sites, blocked by 5,568 (15.6% of those that name it) rogerbot seo 15.6% of 35,595 BLEXBot — named by 20,349 sites, blocked by 18,705 (91.9% of those that name it) BLEXBot seo 91.9% of 20,349 Baiduspider — named by 18,162 sites, blocked by 12,420 (68.4% of those that name it) Baiduspider search 68.4% of 18,162 ia_archiver — named by 12,806 sites, blocked by 9,751 (76.1% of those that name it) ia_archiver archive 76.1% of 12,806 Slurp — named by 11,264 sites, blocked by 2,112 (18.8% of those that name it) Slurp search 18.8% of 11,264

Search crawlers get named often and blocked rarely — being in the index is the point. The interesting middle ground is the SEO-tooling crawlers: AhrefsBot, SemrushBot, MJ12bot and friends provide no traffic in return for the bandwidth they consume, and their block rates look much closer to the AI crawlers' than to Google's.

The archival crawlers sit at the same end. ia_archiver is blocked by 76.1% of the sites that name it — though that number should be read with care: ia_archiver was Alexa's crawler, retired years ago, so much of what we're seeing is fossilised rules nobody has revisited. robots.txt accumulates; it rarely gets pruned.

The pattern is not “sites hate robots.” It's an economic ledger: crawlers that send visitors are welcome, crawlers that only take are not — and the AI crawlers are being sorted, quickly, into the second category.

05 — Rank effects

How far down does the policy go?

Split the corpus by popularity rank. The obvious hypothesis — that elaborate robots.txt policy is a luxury of large sites — is only half right.

File complexity behaves exactly as expected. The top band averages 4,232 bytes against 1,318 at the bottom, a 3.2× spread, and the median file in the tail carries just 8 rules across 1 user-agent group. Sophistication tracks investment, and investment tracks traffic.

1–1k: 23.7% — 122 of 515 sites 23.7% 1–1k 1k–10k: 21.6% — 1,149 of 5,320 sites 21.6% 1k–10k 10k–50k: 19.7% — 4,428 of 22,424 sites 19.7% 10k–50k 50k–100k: 19.9% — 5,912 of 29,652 sites 19.9% 50k–100k 100k–250k: 19.2% — 16,868 of 87,706 sites 19.2% 100k–250k 250k–500k: 16.8% — 25,213 of 150,350 sites 16.8% 250k–500k 500k–1M: 16.0% — 49,353 of 309,079 sites 16.0% 500k–1M
Share of sites blocking at least one AI crawler, by rank band.

AI blocking does not follow that pattern. It runs from 23.7% in the top band to 16.0% at the bottom — a decline of only 1.5× across a thousandfold range of popularity. A site nobody has heard of is nearly as likely to name and refuse an AI crawler as one in the top thousand.

Blocking AI crawlers is the least elitist policy on this page. It reaches essentially the whole web, not just the part of it with a traffic team.

That reach is the finding. Almost every other robots.txt behaviour is a proxy for engineering effort. AI blocking barely is, which suggests it mostly isn't hand-written — it arrives pre-installed, in CMS defaults, WordPress and Shopify plugins, and one-click toggles from CDN providers. A small site inherits the same blocklist a large one deliberates over. The policy diffused across the web faster than the expertise to write it.

1–1k: 7.2% — 37 of 515 sites 7.2% 1–1k 1k–10k: 3.0% — 158 of 5,320 sites 3.0% 1k–10k 10k–50k: 3.4% — 759 of 22,424 sites 3.4% 10k–50k 50k–100k: 3.7% — 1,083 of 29,652 sites 3.7% 50k–100k 100k–250k: 3.8% — 3,310 of 87,706 sites 3.8% 100k–250k 250k–500k: 7.8% — 11,683 of 150,350 sites 7.8% 250k–500k 500k–1M: 10.2% — 31,651 of 309,079 sites 10.2% 500k–1M
Share of sites whose wildcard group is a blanket Disallow: /, by rank band.

Blanket Disallow: / traces a different curve entirely — a U. It is 7.2% at the top, bottoms out at 3.0% around 1k–10k, then climbs back to 10.2% in the deepest band. The two ends are doing opposite things for opposite reasons. Large sites wall off unknown crawlers deliberately, because they have infrastructure to protect. The deep tail is thick with parked domains, staging servers, expired projects and "coming soon" pages — sites that block everything not as policy but as a side effect of never having been launched. The healthy, ordinary web sits in the trough between them.

Rank bandQueriedServed robotsServe rateNames an AI botBlocks an AI botBlanket blockSitemapMean size
1–1k1,00051551.5%36.7%23.7%7.2%69.1%4,232 B
1k–10k9,0005,32059.1%28.5%21.6%3.0%66.7%2,178 B
10k–50k40,00022,42456.1%26.1%19.7%3.4%65.3%1,816 B
50k–100k50,00029,65259.3%24.4%19.9%3.7%62.8%1,463 B
100k–250k150,00087,70658.5%23.5%19.2%3.8%61.2%1,472 B
250k–500k250,000150,35060.1%25.7%16.8%7.8%64.5%1,441 B
500k–1M500,000309,07961.8%25.9%16.0%10.2%63.2%1,318 B

06 — What gets blocked

The anatomy of a disallow

Strip out the blanket Disallow: / and a remarkably consistent vocabulary emerges. The web blocks the same handful of things everywhere.

/wp-admin/: 76,453 /wp-admin/ 76,453 12.6% of sites /admin/: 65,735 /admin/ 65,735 10.9% of sites /admin: 43,512 /admin 43,512 7.2% of sites /checkout: 38,541 /checkout 38,541 6.4% of sites /account: 35,467 /account 35,467 5.9% of sites /search/: 34,885 /search/ 34,885 5.8% of sites /orders: 33,229 /orders 33,229 5.5% of sites /checkouts/: 32,103 /checkouts/ 32,103 5.3% of sites /collections/*%2B*: 31,902 /collections/*%2B* 31,902 5.3% of sites /collections/*+*: 31,894 /collections/*+* 31,894 5.3% of sites /blogs/*+*: 31,859 /blogs/*+* 31,859 5.3% of sites /collections/*%2b*: 31,857 /collections/*%2b* 31,857 5.3% of sites /blogs/*%2B*: 31,850 /blogs/*%2B* 31,850 5.3% of sites /blogs/*%2b*: 31,826 /blogs/*%2b* 31,826 5.3% of sites /collections/*sort_by*: 31,777 /collections/*sort_by* 31,777 5.3% of sites /*/collections/*%2B*: 31,746 /*/collections/*%2B* 31,746 5.2% of sites /*/collections/*sort_by*: 31,740 /*/collections/*sort_by* 31,740 5.2% of sites /*/collections/*%2b*: 31,730 /*/collections/*%2b* 31,730 5.2% of sites /*/blogs/*+*: 31,726 /*/blogs/*+* 31,726 5.2% of sites /*/blogs/*%2B*: 31,722 /*/blogs/*%2B* 31,722 5.2% of sites
Most common disallowed paths, counted once per site. Blanket / and /* rules excluded.

Three motives cover nearly all of it. Infinite spaces — search results, filters, calendars, sort parameters — where a crawler can generate URLs forever and index nothing of value. Private surfaces — admin panels, carts, checkouts, account pages — that leak nothing useful and sometimes leak something sensitive. And machinery — print views, build assets, tracking redirects — that was never meant to be a page.

One entry gives the game away. /wp-admin/ appears in 76,453 files, 12.6% of the whole corpus and the single most disallowed path on the web. That is not a decision anyone made; it is WordPress shipping a default. A large share of the web's robots.txt rules were written once, by a CMS vendor, and inherited wholesale — the same mechanism that put AI blocklists on small sites in the section above.

07 — Who gets addressed

CCBot is the most-addressed crawler on the web

Rank every user-agent token by how many sites name it, set aside the universal *, and the result is the clearest single statement this dataset makes.

The 8 most-named crawlers on the web are, without exception, AI crawlers. CCBot leads at 119,287 sites — 1.7× as many as name Googlebot, which sits at 10th. In under three years, AI crawlers have accumulated more site-specific robots.txt rules than search engines did in twenty-five.

CCBot — named in 119,287 files CCBot 119,287 19.7% Bytespider — named in 112,269 files Bytespider 112,269 18.6% GPTBot — named in 102,412 files GPTBot 102,412 16.9% ClaudeBot — named in 93,580 files ClaudeBot 93,580 15.5% Google-Extended — named in 86,346 files Google-Extended 86,346 14.3% Amazonbot — named in 84,408 files Amazonbot 84,408 14.0% Meta-ExternalAgent — named in 79,302 files Meta-ExternalAgent 79,302 13.1% Applebot-Extended — named in 76,643 files Applebot-Extended 76,643 12.7% AhrefsBot — named in 69,836 files AhrefsBot 69,836 11.5% Googlebot — named in 69,422 files Googlebot 69,422 11.5% MJ12bot — named in 67,310 files MJ12bot 67,310 11.1% SemrushBot — named in 59,584 files SemrushBot 59,584 9.8% Bingbot — named in 56,092 files Bingbot 56,092 9.3% cloudflarebrowserrenderingcrawler — named in 55,100 files cloudflarebrowserrendering… 55,100 9.1% DotBot — named in 53,558 files DotBot 53,558 8.9% PetalBot — named in 49,836 files PetalBot 49,836 8.2% AdsBot-Google — named in 47,955 files AdsBot-Google 47,955 7.9% DataForSeoBot — named in 41,716 files DataForSeoBot 41,716 6.9%
User-agent tokens by number of robots.txt files naming them, excluding the universal * (present in 566,317 files).

The web has spent a quarter century writing rules for crawlers that send traffic. It is now writing far more rules, far faster, for crawlers that don't.

A caveat on reading this chart: naming a crawler is not the same as blocking it, and the tokens here are counted exactly as written. A site that writes User-agent: Applebot has addressed Apple's search crawler, not Applebot-Extended, its separate AI-training opt-out — they are different products and this analysis keeps them apart. Sites frequently conflate the two, which is its own kind of finding.

08 — Hygiene

Structure, size, and sloppiness

robots.txt has no validator, no build step, and no error message. It shows.

MeasureValue
Median file size256 bytes
90th percentile size3,072 bytes
99th percentile size11,264 bytes
Largest file524,288 bytes
Median user-agent groups per file1
Mean rules per file28.1
Files declaring a sitemap383,101 (63.3%)
Sitemap URLs declared in total899,295
Files with zero parseable groups28,178
Files with rules before any User-agent7,399
Files with a BOM or non-breaking space4,420
Files exceeding the 512 KB read cap180

Path syntax tells a similar story. Wildcards (*) and end-anchors ($) were vendor extensions for two decades and only became standard in RFC 9309 in 2022 — yet 34.9% of files use a wildcard and 8.5% use an anchor. A crawler implementing the literal 1994 draft would read those patterns as literal characters and match nothing, quietly ignoring rules the site owner believes are in force. Allow, likewise absent from the original draft, now appears in 57.8% of files.

Non-standard directives are common enough to be worth naming. Crawl-delay is understood by Bing and Yandex but ignored by Google; Host and Clean-param are Yandex extensions; Noindex in robots.txt has never worked and was formally killed by Google in 2019 — yet it persists.

DirectiveOccurrencesStandard?
disallow14,074,433RFC 9309
user-agent4,895,228RFC 9309
allow2,913,372RFC 9309
sitemap899,498RFC 9309
crawl-delay169,609Vendor extension
clean-param75,959Vendor extension
content-signal64,315Non-standard
host25,461Vendor extension
noindex20,243Non-standard
https7,636Non-standard
color4,671Non-standard
display4,298Non-standard

Directives no major crawler implements, seen in the wild: content-signal llm-policy llms display width background height padding margin https.

Crawl-delay values cluster on round numbers: 10s (37,167), 1s (16,104), 5s (11,519), 30s (6,436), 2s (5,503), 20s (4,150).

09 — Extremes

The outliers

A few sites treat robots.txt as a configuration file of record.

DomainRankSizeRules
eerospeedtests.com#396,164524,288 B0
mywaifu.best#24,948480,299 B0
emanuelemiani.it#508,527475,186 B0
quakeworld.nu#457,305475,084 B0
maxgoodell.com#267,993474,993 B0
moe.team#257,165471,782 B0
whenbuff.com#191,333432,678 B0
idealbimbo.it#785,947262,208 B4,739
figma.com#867262,144 B5,474
boe.es#4,671262,144 B6,555
trademe.co.nz#13,740262,144 B6,031
xuexi.cn#14,140262,144 B3,319
mercer.com#18,514262,144 B3,324
fucolle.com#22,571262,144 B9,045
conference-board.org#31,596262,144 B1,274

And the sites naming the most AI crawlers individually — each of these maintains an explicit, hand-curated blocklist rather than a blanket rule:

DomainRankAI crawlers blocked
theconversation.com#99929
furaffinity.net#1,22929
themoviedb.org#2,05529
tmdb.org#2,68329
imageshack.us#2,90129
metro.co.uk#3,08829
newscientist.com#2,67029
politico.eu#3,52429
thestar.com#4,02629
imageshack.com#4,82729
gentoo.org#5,60829
refinery29.com#7,62729
ylilauta.org#9,87629
inews.co.uk#10,12029
augsburger-allgemeine.de#11,04429

Method & caveats

Domains come from the Tranco top-1M daily list, a research ranking that averages several commercial popularity lists to resist manipulation. We took the top 1,000,000 and issued one GET to https://<domain>/robots.txt, falling back to HTTP on connection failure and retrying with a www. prefix where the apex domain did not resolve. Redirects were followed; responses were capped at 512 KB; the crawler identified itself honestly and did not retry aggressively.

Parsing follows RFC 9309: consecutive User-agent lines share the rule block that follows, and the most specific matching agent group wins over *. “Blocked” means the group that applies to that agent contains Disallow: / with no Allow carve-out. Path counts are deduplicated per site, so a file repeating Disallow: / across forty bot groups counts once.

Caveats worth stating plainly. A popularity list is not a list of websites — infrastructure hostnames inflate the DNS-failure bucket. Sites behind bot protection may serve us a different robots.txt, or a 403, than they serve Googlebot. Absence of a rule is not permission, and presence of one is not enforcement: robots.txt is a request, honored voluntarily, and the crawlers most likely to ignore it are exactly the ones these files increasingly name.

Dataset: data/robots-1m-dataset.ndjson.gz · Aggregates: data/analysis.json · Generated 2026-08-07T19:52:13.527Z