Every month Desvela crawls the top 100,000 domains for the four surfaces a visiting AI agent can actually read: ai-catalog.json (Google's ARD standard), llms.txt, agents.md, and the AI-bot directives in robots.txt. This is what we found.
GPTBot block it entirely| Tranco tier | Checked | ai-catalog.json | llms.txt | agents.md |
|---|---|---|---|---|
| top 1K | 1,000 | 2 (0.2%) | 69 (6.9%) | 2 (0.2%) |
| 1K–10K | 9,000 | 19 (0.2%) | 552 (6.1%) | 21 (0.2%) |
| 10K–100K | 90,000 | 126 (0.1%) | 5447 (6.1%) | 901 (1.0%) |
Domains that name each bot in robots.txt, and how many of those shut it out completely (Disallow: /). Before your agent touches a domain, this is the etiquette it should know.
| Bot | Named in robots.txt | Fully blocked |
|---|---|---|
GPTBot | 10,288 | 8,174 (79.5%) |
ClaudeBot | 9,156 | 7,277 (79.5%) |
CCBot | 8,900 | 7,900 (88.8%) |
Google-Extended | 8,315 | 6,722 (80.8%) |
Bytespider | 8,168 | 7,568 (92.7%) |
meta-externalagent | 7,381 | 6,483 (87.8%) |
Applebot-Extended | 7,135 | 6,352 (89.0%) |
PerplexityBot | 4,191 | 2,252 (53.7%) |
anthropic-ai | 3,497 | 2,600 (74.3%) |
cohere-ai | 2,712 | 2,175 (80.2%) |
Claude-Web | 2,643 | 2,043 (77.3%) |
The ARD spec added a DNS path in July 2026: publish an SVCB record at _index._agents.<domain> and an agent can find your registry without fetching anything over HTTP. We probed 99,048 domains and 15 publish one, which is 0.015% or about 1 in 6,603.
ai-catalog.json (valid, with entries): huggingface.co #1,263 · hostinger.com #1,273 · zapier.com #2,834 · airtable.com #2,925 · gtmetrix.com #11,362 · opus.pro #24,409 · roboflow.com #29,085 · fundraiseup.com #32,223 · clickhouse.com #32,343 · datarobot.com #55,938 · idescat.cat #63,234 · guruwalk.com #81,321
llms.txt: cloudflare.com #3 · github.com #31 · wordpress.org #43 · digicert.com #45 · adobe.com #67 · opera.com #90 · samsung.com #94 · wordpress.com #101 · dropbox.com #136 · kaspersky.com #139 · gravatar.com #141 · paypal.com #163
agents.md: checkpoint.com #328 · paloaltonetworks.com #491 · crazyegg.com #2,894 · razorpay.com #3,649 · nazwa.pl #4,831 · eufylife.com #4,913 · kvant-telecom.ru #5,082 · laravel-news.com #5,219 · guinnessworldrecords.com #5,869 · jbhifi.com.au #6,537 · wyze.com #6,627 · mattel.com #8,038
Same four files, same rules, same canary request. You get a grade against the numbers above, and the specific file you are missing. Free, no account, about ten seconds.
Each edition stays at its own URL after it is replaced, so a number you cited last month is still there to check.
August 2026 (this one) · July 2026
Plain GET requests (no rendering) against /.well-known/ai-catalog.json, /llms.txt, /agents.md and /robots.txt on each apex domain of the Tranco top-100K, paced at ~4 req/s globally, honoring robots.txt, as DesvelaBot/0.1 (bot policy). HTML responses to text paths count as absent. Every publisher is re-probed with a canary request to a nonexistent path: servers that answer 200 to anything are excluded as catch-alls, including some seemingly legitimate publishers. Undercounting beats inflating.
One email per month with the fresh numbers and what changed: new publishers, withdrawals, bot-blocking shifts. No product spam, unsubscribe by replying.