Dataset & Analysis
What the web's busiest sites refuse to let you crawl
We asked the top 1,000,000 domains on the internet for their
robots.txt, parsed every rule, and counted what they allow, what they block,
and — increasingly — who they block.
01 — Coverage
Asking 1,000,000 sites a simple question
A request for /robots.txt is the most basic thing a crawler does.
Only 60.5% of the top 1,000,000 domains answered it with an actual
robots file.
The rest split into instructive failure modes. Some domains in a popularity ranking aren't
websites at all — CDN edges, nameserver infrastructure, and tracking hosts inflate the DNS
failures. Others actively refuse: 74,194 returned an HTTP error,
most of them 403 from bot-protection layers that block the very file designed to
tell bots what to do. And 47,958 served an HTML page with a
200 status — a soft 404 that a naive parser would happily read as a robots file
with zero rules.
- Served robots.txt 605,046 60.5%
- Network / DNS failure 164,682 16.5%
- HTTP error (403, 5xx…) 74,194 7.4%
- 404 / 410 — no file 100,591 10.1%
- Unfollowable redirect 69 0.0%
- HTML page, not robots 47,958 4.8%
- Empty file 7,460 0.7%
/robots.txt per domain, HTTPS with an HTTP fallback and a www. retry for domains whose apex does not resolve.text/plain-shaped content. Treat those as valid and
you conclude they permit everything — the exact opposite of what many of them intend.02 — Default posture
Most of the web is open, conditionally
For the generic User-agent: * group — the rules that apply to a
crawler nobody has heard of — the dominant answer is yes, but not there.
- Open — no restrictions 173,246 28.6%
- Partial — some paths blocked 344,390 56.9%
- Blocked — Disallow: / 48,681 8.0%
- No wildcard group 38,729 6.4%
Disallow: / with no Allow carve-out.Only 48,681 sites (8.0%) slam the door on unknown crawlers entirely. The typical file is a scalpel, not a wall: a median of 8 rules across 1 user-agent group, carving out the parts of a site that are expensive to serve or worthless to index.
03 — The AI question
The new names on the blocklist
The most significant change to robots.txt in twenty years is not a new directive. It's a new class of user-agent — and site owners are naming them explicitly.
154,160 sites (25.5% of those serving a robots file) name at least one AI or LLM-training crawler by name. Of those, the overwhelming majority are not writing rules to welcome them: 103,045 sites block at least one outright.
When a site bothers to mention an AI crawler at all, the mention is usually a refusal — a block rate far above anything the traditional search crawlers see.
| Crawler | Operator | Sites naming it | % of corpus | Blocked by | Block rate |
|---|---|---|---|---|---|
CCBot | Common Crawl | 119,287 | 19.72% | 79,830 | 66.9% |
Bytespider | ByteDance | 112,269 | 18.56% | 76,212 | 67.9% |
GPTBot | OpenAI | 102,412 | 16.93% | 80,957 | 79.1% |
ClaudeBot | Anthropic | 93,580 | 15.47% | 74,561 | 79.7% |
Google-Extended | 86,346 | 14.27% | 69,755 | 80.8% | |
Amazonbot | Amazon | 84,408 | 13.95% | 74,037 | 87.7% |
Meta-ExternalAgent | Meta | 79,302 | 13.11% | 67,078 | 84.6% |
Applebot-Extended | Apple | 76,643 | 12.67% | 67,021 | 87.4% |
PetalBot | Huawei | 49,836 | 8.24% | 17,474 | 35.1% |
ChatGPT-User | OpenAI | 32,649 | 5.4% | 15,967 | 48.9% |
PerplexityBot | Perplexity | 31,283 | 5.17% | 13,070 | 41.8% |
anthropic-ai | Anthropic | 26,444 | 4.37% | 16,114 | 60.9% |
OAI-SearchBot | OpenAI | 21,643 | 3.58% | 6,409 | 29.6% |
Claude-Web | Anthropic | 19,844 | 3.28% | 12,607 | 63.5% |
cohere-ai | Cohere | 19,579 | 3.24% | 12,756 | 65.2% |
FacebookBot | Meta | 17,050 | 2.82% | 10,187 | 59.7% |
YouBot | You.com | 16,709 | 2.76% | 10,570 | 63.3% |
omgili | Webz.io | 16,372 | 2.71% | 12,666 | 77.4% |
omgilibot | Webz.io | 16,048 | 2.65% | 12,228 | 76.2% |
ImagesiftBot | Imagesift | 15,147 | 2.5% | 12,241 | 80.8% |
Diffbot | Diffbot | 14,239 | 2.35% | 10,451 | 73.4% |
Timpibot | Timpi | 10,189 | 1.68% | 7,558 | 74.2% |
04 — For contrast
Nobody blocks Googlebot
Run the same measurement against the crawlers that have been around for decades and the picture inverts.
Search crawlers get named often and blocked rarely — being in the index is the point. The
interesting middle ground is the SEO-tooling crawlers: AhrefsBot,
SemrushBot, MJ12bot and friends provide no traffic in return for the
bandwidth they consume, and their block rates look much closer to the AI crawlers' than to
Google's.
The archival crawlers sit at the same end. ia_archiver is blocked by
76.1% of the sites
that name it — though that number should be read with care: ia_archiver was Alexa's
crawler, retired years ago, so much of what we're seeing is fossilised rules nobody has revisited.
robots.txt accumulates; it rarely gets pruned.
05 — Rank effects
How far down does the policy go?
Split the corpus by popularity rank. The obvious hypothesis — that elaborate robots.txt policy is a luxury of large sites — is only half right.
File complexity behaves exactly as expected. The top band averages 4,232 bytes against 1,318 at the bottom, a 3.2× spread, and the median file in the tail carries just 8 rules across 1 user-agent group. Sophistication tracks investment, and investment tracks traffic.
AI blocking does not follow that pattern. It runs from 23.7% in the top band to 16.0% at the bottom — a decline of only 1.5× across a thousandfold range of popularity. A site nobody has heard of is nearly as likely to name and refuse an AI crawler as one in the top thousand.
Blocking AI crawlers is the least elitist policy on this page. It reaches essentially the whole web, not just the part of it with a traffic team.
That reach is the finding. Almost every other robots.txt behaviour is a proxy for engineering effort. AI blocking barely is, which suggests it mostly isn't hand-written — it arrives pre-installed, in CMS defaults, WordPress and Shopify plugins, and one-click toggles from CDN providers. A small site inherits the same blocklist a large one deliberates over. The policy diffused across the web faster than the expertise to write it.
Disallow: /, by rank band.Blanket Disallow: / traces a different curve entirely — a U. It is
7.2% at the top, bottoms out at 3.0% around
1k–10k, then climbs back to 10.2%
in the deepest band. The two ends are doing opposite things for opposite reasons. Large sites
wall off unknown crawlers deliberately, because they have infrastructure to protect. The deep
tail is thick with parked domains, staging servers, expired projects and "coming soon" pages —
sites that block everything not as policy but as a side effect of never having been launched.
The healthy, ordinary web sits in the trough between them.
| Rank band | Queried | Served robots | Serve rate | Names an AI bot | Blocks an AI bot | Blanket block | Sitemap | Mean size |
|---|---|---|---|---|---|---|---|---|
1–1k | 1,000 | 515 | 51.5% | 36.7% | 23.7% | 7.2% | 69.1% | 4,232 B |
1k–10k | 9,000 | 5,320 | 59.1% | 28.5% | 21.6% | 3.0% | 66.7% | 2,178 B |
10k–50k | 40,000 | 22,424 | 56.1% | 26.1% | 19.7% | 3.4% | 65.3% | 1,816 B |
50k–100k | 50,000 | 29,652 | 59.3% | 24.4% | 19.9% | 3.7% | 62.8% | 1,463 B |
100k–250k | 150,000 | 87,706 | 58.5% | 23.5% | 19.2% | 3.8% | 61.2% | 1,472 B |
250k–500k | 250,000 | 150,350 | 60.1% | 25.7% | 16.8% | 7.8% | 64.5% | 1,441 B |
500k–1M | 500,000 | 309,079 | 61.8% | 25.9% | 16.0% | 10.2% | 63.2% | 1,318 B |
06 — What gets blocked
The anatomy of a disallow
Strip out the blanket Disallow: / and a remarkably consistent
vocabulary emerges. The web blocks the same handful of things everywhere.
/ and /* rules excluded.Three motives cover nearly all of it. Infinite spaces — search results, filters, calendars, sort parameters — where a crawler can generate URLs forever and index nothing of value. Private surfaces — admin panels, carts, checkouts, account pages — that leak nothing useful and sometimes leak something sensitive. And machinery — print views, build assets, tracking redirects — that was never meant to be a page.
One entry gives the game away. /wp-admin/ appears in
76,453 files, 12.6%
of the whole corpus and the single most disallowed path on the web.
That is not a decision anyone made; it is WordPress shipping a default. A large share of the
web's robots.txt rules were written once, by a CMS vendor, and inherited wholesale — the same
mechanism that put AI blocklists on small sites in the section above.
07 — Who gets addressed
CCBot is the most-addressed crawler on the web
Rank every user-agent token by how many sites name it, set aside the universal
*, and the result is the clearest single statement this dataset makes.
The 8 most-named crawlers on the web are, without exception, AI
crawlers. CCBot leads at 119,287
sites — 1.7× as many as name
Googlebot, which sits at 10th.
In under three years, AI crawlers have accumulated more site-specific robots.txt rules than
search engines did in twenty-five.
* (present in 566,317 files).The web has spent a quarter century writing rules for crawlers that send traffic. It is now writing far more rules, far faster, for crawlers that don't.
A caveat on reading this chart: naming a crawler is not the same as blocking it, and the tokens
here are counted exactly as written. A site that writes User-agent: Applebot has
addressed Apple's search crawler, not Applebot-Extended, its separate AI-training
opt-out — they are different products and this analysis keeps them apart. Sites frequently
conflate the two, which is its own kind of finding.
08 — Hygiene
Structure, size, and sloppiness
robots.txt has no validator, no build step, and no error message. It shows.
| Measure | Value |
|---|---|
| Median file size | 256 bytes |
| 90th percentile size | 3,072 bytes |
| 99th percentile size | 11,264 bytes |
| Largest file | 524,288 bytes |
| Median user-agent groups per file | 1 |
| Mean rules per file | 28.1 |
| Files declaring a sitemap | 383,101 (63.3%) |
| Sitemap URLs declared in total | 899,295 |
| Files with zero parseable groups | 28,178 |
Files with rules before any User-agent | 7,399 |
| Files with a BOM or non-breaking space | 4,420 |
| Files exceeding the 512 KB read cap | 180 |
Path syntax tells a similar story. Wildcards (*) and end-anchors
($) were vendor extensions for two decades and only became standard in RFC 9309 in
2022 — yet 34.9% of files use a wildcard and
8.5% use an anchor. A crawler implementing the literal 1994
draft would read those patterns as literal characters and match nothing, quietly ignoring rules
the site owner believes are in force. Allow, likewise absent from the original draft,
now appears in 57.8% of files.
Non-standard directives are common enough to be worth naming. Crawl-delay is
understood by Bing and Yandex but ignored by Google; Host and
Clean-param are Yandex extensions; Noindex in robots.txt has never
worked and was formally killed by Google in 2019 — yet it persists.
| Directive | Occurrences | Standard? |
|---|---|---|
disallow | 14,074,433 | RFC 9309 |
user-agent | 4,895,228 | RFC 9309 |
allow | 2,913,372 | RFC 9309 |
sitemap | 899,498 | RFC 9309 |
crawl-delay | 169,609 | Vendor extension |
clean-param | 75,959 | Vendor extension |
content-signal | 64,315 | Non-standard |
host | 25,461 | Vendor extension |
noindex | 20,243 | Non-standard |
https | 7,636 | Non-standard |
color | 4,671 | Non-standard |
display | 4,298 | Non-standard |
Directives no major crawler implements, seen in the wild:
content-signal llm-policy llms display width background height padding margin https.
Crawl-delay values cluster on round numbers:
10s (37,167), 1s (16,104), 5s (11,519), 30s (6,436), 2s (5,503), 20s (4,150).
09 — Extremes
The outliers
A few sites treat robots.txt as a configuration file of record.
| Domain | Rank | Size | Rules |
|---|---|---|---|
eerospeedtests.com | #396,164 | 524,288 B | 0 |
mywaifu.best | #24,948 | 480,299 B | 0 |
emanuelemiani.it | #508,527 | 475,186 B | 0 |
quakeworld.nu | #457,305 | 475,084 B | 0 |
maxgoodell.com | #267,993 | 474,993 B | 0 |
moe.team | #257,165 | 471,782 B | 0 |
whenbuff.com | #191,333 | 432,678 B | 0 |
idealbimbo.it | #785,947 | 262,208 B | 4,739 |
figma.com | #867 | 262,144 B | 5,474 |
boe.es | #4,671 | 262,144 B | 6,555 |
trademe.co.nz | #13,740 | 262,144 B | 6,031 |
xuexi.cn | #14,140 | 262,144 B | 3,319 |
mercer.com | #18,514 | 262,144 B | 3,324 |
fucolle.com | #22,571 | 262,144 B | 9,045 |
conference-board.org | #31,596 | 262,144 B | 1,274 |
And the sites naming the most AI crawlers individually — each of these maintains an explicit, hand-curated blocklist rather than a blanket rule:
| Domain | Rank | AI crawlers blocked |
|---|---|---|
theconversation.com | #999 | 29 |
furaffinity.net | #1,229 | 29 |
themoviedb.org | #2,055 | 29 |
tmdb.org | #2,683 | 29 |
imageshack.us | #2,901 | 29 |
metro.co.uk | #3,088 | 29 |
newscientist.com | #2,670 | 29 |
politico.eu | #3,524 | 29 |
thestar.com | #4,026 | 29 |
imageshack.com | #4,827 | 29 |
gentoo.org | #5,608 | 29 |
refinery29.com | #7,627 | 29 |
ylilauta.org | #9,876 | 29 |
inews.co.uk | #10,120 | 29 |
augsburger-allgemeine.de | #11,044 | 29 |
Method & caveats
Domains come from the Tranco top-1M daily list, a
research ranking that averages several commercial popularity lists to resist manipulation. We
took the top 1,000,000 and issued one GET to https://<domain>/robots.txt,
falling back to HTTP on connection failure and retrying with a www. prefix where the
apex domain did not resolve. Redirects were followed; responses were capped at 512 KB;
the crawler identified itself honestly and did not retry aggressively.
Parsing follows RFC 9309: consecutive User-agent lines share the rule block that
follows, and the most specific matching agent group wins over *. “Blocked” means the
group that applies to that agent contains Disallow: / with no Allow
carve-out. Path counts are deduplicated per site, so a file repeating
Disallow: / across forty bot groups counts once.
Caveats worth stating plainly. A popularity list is not a list of websites — infrastructure
hostnames inflate the DNS-failure bucket. Sites behind bot protection may serve us a different
robots.txt, or a 403, than they serve Googlebot. Absence of a rule is not permission,
and presence of one is not enforcement: robots.txt is a request, honored voluntarily, and the
crawlers most likely to ignore it are exactly the ones these files increasingly name.
Dataset: data/robots-1m-dataset.ndjson.gz ·
Aggregates: data/analysis.json · Generated 2026-08-07T19:52:13.527Z