Every month Desvela crawls the top 100,000 domains for the files an AI agent could read. Then it asks each of them for a path that cannot exist. A server that answers that answers anything, so everything it appears to publish is an artefact of the server, not a decision by its owner. This month that removed 93% of the domains that appear to publish Google's ARD manifest.
This share is not comparable with earlier editions. The method changed this month (v2.0), and it changed in the direction of finding more of them: the canary now probes /.well-known/ as well as the root, and servers that answer only under /.well-known/ were previously counted as publishers. The web did not get worse between editions; the instrument got better at seeing it. Details in the manifest below.
mcp.json — down from 794 that answer on that pathagent-card.json — down from 754 that answer on that pathai-catalog.json — down from 763 that answer on that pathard.json — down from 727 that answer on that path| Tranco tier | Checked | ai-catalog.json | llms.txt | agents.md | agent-card.json | mcp.json | ard.json |
|---|---|---|---|---|---|---|---|
| top 1K | 1,000 | 0 (0.0%) | 75 (7.5%) | 4 (0.4%) | 2 (0.2%) | 3 (0.3%) | 0 (0.0%) |
| 1K–10K | 9,000 | 6 (0.1%) | 555 (6.2%) | 18 (0.2%) | 6 (0.1%) | 9 (0.1%) | 0 (0.0%) |
| 10K–100K | 89,999 | 31 (0.0%) | 5724 (6.4%) | 949 (1.1%) | 45 (0.1%) | 72 (0.1%) | 4 (0.0%) |
Domains that name each bot in robots.txt, and how many of those shut it out completely (Disallow: /). Before your agent touches a domain, this is the etiquette it should know.
| Bot | Named in robots.txt | Fully blocked |
|---|---|---|
GPTBot | 10,540 | 8,307 (78.8%) |
ClaudeBot | 9,527 | 7,508 (78.8%) |
CCBot | 9,180 | 8,101 (88.2%) |
Google-Extended | 8,660 | 6,914 (79.8%) |
Bytespider | 8,474 | 7,814 (92.2%) |
meta-externalagent | 7,817 | 6,818 (87.2%) |
Applebot-Extended | 7,527 | 6,646 (88.3%) |
PerplexityBot | 4,291 | 2,157 (50.3%) |
ChatGPT-User | 4,193 | 2,147 (51.2%) |
anthropic-ai | 3,442 | 2,489 (72.3%) |
cohere-ai | 2,691 | 2,094 (77.8%) |
Claude-Web | 2,649 | 1,980 (74.7%) |
Perplexity-User | 1,579 | 775 (49.1%) |
Claude-User | 1,502 | 747 (49.7%) |
DuckAssistBot | 1,297 | 922 (71.1%) |
meta-externalfetcher | 1,297 | 947 (73.0%) |
MistralAI-User | 843 | 621 (73.7%) |
Gemini-Deep-Research | 443 | 322 (72.7%) |
NovaAct | 364 | 334 (91.8%) |
Google-NotebookLM | 340 | 247 (72.6%) |
Devin | 325 | 309 (95.1%) |
GoogleAgent-Mariner | 292 | 252 (86.3%) |
AmazonBuyForMe | 249 | 241 (96.8%) |
Manus-User | 230 | 218 (94.8%) |
TwinAgent | 205 | 201 (98.0%) |
Google-Agent | 186 | 148 (79.6%) |
Kagi-Fetcher | 78 | 74 (94.9%) |
Claude-Code | 62 | 48 (77.4%) |
Trae | 49 | 48 (98.0%) |
OpenCode | 49 | 47 (95.9%) |
Kimi-User | 46 | 42 (91.3%) |
GoogleAgent-URLContext | 40 | 37 (92.5%) |
Google-Gemini-CLI | 35 | 32 (91.4%) |
Cursor | 22 | 20 (90.9%) |
ChatGPT-Agent | 8 | 4 (50.0%) |
The table above is what domains declare in robots.txt. This one is what happened when we sent each AI crawler's real user-agent at the homepage of the Tranco top-1,000 and compared the response to a browser's, from the same connection. The headline is not the blocking: 413 of 999 domains (41.3%) never answered our browser-labelled control either — they refuse any plain HTTP client, whoever it claims to be.
| Crawler | Blocked (4xx) | Throttled (429) | Degraded (<60% of bytes) | Did not get the page |
|---|---|---|---|---|
ClaudeBot | 74 | 12 | 8 | 16.0% |
GPTBot | 73 | 4 | 9 | 14.7% |
CCBot | 69 | 7 | 9 | 14.5% |
PerplexityBot | 63 | 3 | 11 | 13.1% |
OAI-SearchBot | 52 | 3 | 8 | 10.8% |
Rates are over the 586 measurable domains — the ones that answered a plain client at all, which are by definition the most permissive of the cohort. The true rate across the full 1,000 is higher and cannot be measured this way. A 429 is counted apart from a block because it says "too fast", not "not you". Swept 2026-09-01 from a residential connection; the vantage is part of the method — datacenter IPs get refused far more often.
The ARD spec added a DNS path in July 2026: publish an SVCB record at _index._agents.<domain> and an agent can find your registry without fetching anything over HTTP. We probed 98,330 domains and 24 publish one, which is 0.024% or about 1 in 4,097.
ai-catalog.json (valid, with entries): huggingface.co #1,176 · hostinger.com #1,504 · padlet.com #2,759 · zapier.com #2,906 · airtable.com #2,949 · rudderstack.com #3,057 · nextjs.org #10,756 · gtmetrix.com #11,507 · speedof.me #13,035 · apify.com #13,443 · bestprice.gr #15,352 · railway.com #18,185
llms.txt: cloudflare.com #2 · github.com #29 · wordpress.org #47 · adobe.com #68 · opera.com #92 · samsung.com #94 · wordpress.com #98 · capgemini.com #101 · kaspersky.com #140 · dropbox.com #142 · gravatar.com #157 · paypal.com #166
agents.md: shopify.com #175 · checkpoint.com #350 · paloaltonetworks.com #410 · shop.app #637 · crazyegg.com #2,683 · razorpay.com #3,285 · nazwa.pl #4,223 · openweathermap.org #4,241 · nanit.com #4,386 · eufylife.com #4,898 · moovitapp.com #5,093 · laravel-news.com #5,176
Same four files, same rules, same canary request. You get a grade against the numbers above, and the specific file you are missing. Free, no account, about ten seconds.
Every number on this page is anchored to one crawl, one universe and one method version. These signals measure adoption and access — a domain publishing a file, a bot being named or blocked — and they do not demonstrate citations, recommendations or visibility inside any AI system.
| Field | Value |
|---|---|
| Edition | 2026-09 |
| Data collected | 2026-09-02 |
| Published | 2026-09-03 |
| Corrected | — |
| Methodology | v2.0 |
| Coverage | 99,999 of 100,000 domains checked |
| Universe | Tranco top-100000 (list 2026-09-01) |
| Own domains in scope | none |
JSON (manifest + metrics) · CSV (by tier)
Methodology, in full: how we count.
Each edition stays at its own URL after it is replaced, so a number you cited last month is still there to check. Only editions with complete or explicitly declared coverage are published.
September 2026 (this one) · August 2026 · July 2026
Plain GET requests (no rendering) against /.well-known/ai-catalog.json, /llms.txt, /agents.md and /robots.txt on each apex domain of the Tranco top-100K, paced at ~4 req/s globally, honoring robots.txt, as DesvelaBot/0.1 (bot policy). HTML responses to text paths count as absent. Every publisher is re-probed with a canary request to a nonexistent path: servers that answer 200 to anything are excluded as catch-alls, including some seemingly legitimate publishers. Undercounting beats inflating.
Read the full methodology → (universe, exclusions, coverage rules and what these numbers do not claim).
One email per month with the fresh numbers and what changed: new publishers, withdrawals, bot-blocking shifts. No product spam. Unsubscribe in one click from any email. We email you a confirmation link first, and nothing is sent until you click it.