State of the agent-readable web · September 2026

Most of what looks like
an agent-readable web is not there.

Every month Desvela crawls the top 100,000 domains for the files an AI agent could read. Then it asks each of them for a path that cannot exist. A server that answers that answers anything, so everything it appears to publish is an artefact of the server, not a decision by its owner. This month that removed 93% of the domains that appear to publish Google's ARD manifest.

This share is not comparable with earlier editions. The method changed this month (v2.0), and it changed in the direction of finding more of them: the canary now probes /.well-known/ as well as the root, and servers that answer only under /.well-known/ were previously counted as publishers. The web did not get worse between editions; the instrument got better at seeing it. Details in the manifest below.

99,999
domains checked this pass
84
real mcp.json — down from 794 that answer on that path
53
real agent-card.json — down from 754 that answer on that path
50
real ai-catalog.json — down from 763 that answer on that path
22
real ard.json — down from 727 that answer on that path

See where your own domain lands →

Adoption

The four surfaces we have measured longest, by ranking tier.

Tranco tierCheckedai-catalog.jsonllms.txtagents.mdagent-card.jsonmcp.jsonard.json
top 1K1,000 0 (0.0%) 75 (7.5%) 4 (0.4%) 2 (0.2%) 3 (0.3%) 0 (0.0%)
1K–10K9,000 6 (0.1%) 555 (6.2%) 18 (0.2%) 6 (0.1%) 9 (0.1%) 0 (0.0%)
10K–100K89,999 31 (0.0%) 5724 (6.4%) 949 (1.1%) 45 (0.1%) 72 (0.1%) 4 (0.0%)
ai-catalog.json reality check: 50 files respond — 10 are invalid JSON, and only 37 are valid manifests with at least one entry: huggingface.co, hostinger.com, padlet.com, zapier.com, airtable.com, rudderstack.com, nextjs.org, gtmetrix.com, speedof.me, apify.com, bestprice.gr, railway.com, mentimeter.com, bird.com, bizzabo.com. Google published the ARD spec in May 2026; adoption is a rounding error.
The door policy

Who's blocking the AI bots.

Domains that name each bot in robots.txt, and how many of those shut it out completely (Disallow: /). Before your agent touches a domain, this is the etiquette it should know.

BotNamed in robots.txtFully blocked
GPTBot10,540 8,307 (78.8%)
ClaudeBot9,527 7,508 (78.8%)
CCBot9,180 8,101 (88.2%)
Google-Extended8,660 6,914 (79.8%)
Bytespider8,474 7,814 (92.2%)
meta-externalagent7,817 6,818 (87.2%)
Applebot-Extended7,527 6,646 (88.3%)
PerplexityBot4,291 2,157 (50.3%)
ChatGPT-User4,193 2,147 (51.2%)
anthropic-ai3,442 2,489 (72.3%)
cohere-ai2,691 2,094 (77.8%)
Claude-Web2,649 1,980 (74.7%)
Perplexity-User1,579 775 (49.1%)
Claude-User1,502 747 (49.7%)
DuckAssistBot1,297 922 (71.1%)
meta-externalfetcher1,297 947 (73.0%)
MistralAI-User843 621 (73.7%)
Gemini-Deep-Research443 322 (72.7%)
NovaAct364 334 (91.8%)
Google-NotebookLM340 247 (72.6%)
Devin325 309 (95.1%)
GoogleAgent-Mariner292 252 (86.3%)
AmazonBuyForMe249 241 (96.8%)
Manus-User230 218 (94.8%)
TwinAgent205 201 (98.0%)
Google-Agent186 148 (79.6%)
Kagi-Fetcher78 74 (94.9%)
Claude-Code62 48 (77.4%)
Trae49 48 (98.0%)
OpenCode49 47 (95.9%)
Kimi-User46 42 (91.3%)
GoogleAgent-URLContext40 37 (92.5%)
Google-Gemini-CLI35 32 (91.4%)
Cursor22 20 (90.9%)
ChatGPT-Agent8 4 (50.0%)
The door policy, observed

What the servers actually do when the crawler shows up.

The table above is what domains declare in robots.txt. This one is what happened when we sent each AI crawler's real user-agent at the homepage of the Tranco top-1,000 and compared the response to a browser's, from the same connection. The headline is not the blocking: 413 of 999 domains (41.3%) never answered our browser-labelled control either — they refuse any plain HTTP client, whoever it claims to be.

CrawlerBlocked (4xx)Throttled (429)Degraded (<60% of bytes)Did not get the page
ClaudeBot74128 16.0%
GPTBot7349 14.7%
CCBot6979 14.5%
PerplexityBot63311 13.1%
OAI-SearchBot5238 10.8%

Rates are over the 586 measurable domains — the ones that answered a plain client at all, which are by definition the most permissive of the cohort. The true rate across the full 1,000 is higher and cannot be measured this way. A 429 is counted apart from a block because it says "too fast", not "not you". Swept 2026-09-01 from a residential connection; the vantage is part of the method — datacenter IPs get refused far more often.

Third-party context, not our measurement: ChatGPT holds 53.9% of assistant web visits and Gemini 27.9% (Similarweb, May 2026, sessions on each assistant's own domain) — yet by referral share arriving at websites, Statcounter (August 2026, its ~1.5M-site tracking network) puts ChatGPT at 79.4% and Gemini at 10.9%, with Perplexity at 4.3% on a 1.3% visit share. The two sources measure different questions and disagree 3× — quote either alone and you have picked a side without saying so. Read together with the table above: the crawler treated best (OAI-SearchBot) belongs to the assistant that sends the most traffic back.
DNS discovery

Almost nobody publishes the DNS record.

The ARD spec added a DNS path in July 2026: publish an SVCB record at _index._agents.<domain> and an agent can find your registry without fetching anything over HTTP. We probed 98,330 domains and 24 publish one, which is 0.024% or about 1 in 4,097.

Every positive is checked against a wildcard canary: we ask the same zone for a name nobody could have published, and if that answers too the positive is discarded. Without that check the number would be higher and wrong.
Notable publishers

Who has set the table.

ai-catalog.json (valid, with entries): huggingface.co #1,176 · hostinger.com #1,504 · padlet.com #2,759 · zapier.com #2,906 · airtable.com #2,949 · rudderstack.com #3,057 · nextjs.org #10,756 · gtmetrix.com #11,507 · speedof.me #13,035 · apify.com #13,443 · bestprice.gr #15,352 · railway.com #18,185

llms.txt: cloudflare.com #2 · github.com #29 · wordpress.org #47 · adobe.com #68 · opera.com #92 · samsung.com #94 · wordpress.com #98 · capgemini.com #101 · kaspersky.com #140 · dropbox.com #142 · gravatar.com #157 · paypal.com #166

agents.md: shopify.com #175 · checkpoint.com #350 · paloaltonetworks.com #410 · shop.app #637 · crazyegg.com #2,683 · razorpay.com #3,285 · nazwa.pl #4,223 · openweathermap.org #4,241 · nanit.com #4,386 · eufylife.com #4,898 · moovitapp.com #5,093 · laravel-news.com #5,176

Your turn

Now check your own domain.

Same four files, same rules, same canary request. You get a grade against the numbers above, and the specific file you are missing. Free, no account, about ten seconds.

Edition manifest

What this edition is.

Every number on this page is anchored to one crawl, one universe and one method version. These signals measure adoption and access — a domain publishing a file, a bot being named or blocked — and they do not demonstrate citations, recommendations or visibility inside any AI system.

FieldValue
Edition2026-09
Data collected2026-09-02
Published2026-09-03
Corrected
Methodologyv2.0
Coverage99,999 of 100,000 domains checked
UniverseTranco top-100000 (list 2026-09-01)
Own domains in scopenone
Methodology 2.0. The canary that discards catch-all servers now probes two directories, the root and /.well-known/, instead of the root alone. Servers that 404 a nonsense path at the root while answering 200 to anything under /.well-known/ were previously counted as publishers of every well-known surface. This is a measurement change, not a change in the web: the catch-all rates below are not comparable with editions 2026-07 and 2026-08, which under-detected them.
Methodology 2.0. A manifest surface (ai-catalog.json, ard.json) now requires a non-empty `entries` array to count. A 200 carrying JSON without entries is a server answering, not a domain publishing. This was already required for ai-catalog.json and now applies to every manifest surface.
Three surfaces appear in the census for the first time: agent-card.json and mcp.json (measured since 2026-08-19) and ard.json (since 2026-08-27, one day after ARD v0.91 moved the canonical path). A zero in an earlier edition means NOT MEASURED, not 'nobody published'.
The four ard.json publishers were each verified by hand against the live path on 2026-09-03, with a canary probe per domain: brandfetch.com, railway.com, idescat.cat and bizzabo.com.
The pass ran across a deployment: domains 0-39,999 were crawled with one build and 40,000-99,999 with the next. The two builds are behaviourally equivalent for every surface measured here — the changes add metadata and a two-domain opt-out list, neither of which affects present/absent. Recorded in the snapshot's .meta.json.

JSON (manifest + metrics) · CSV (by tier)

Methodology, in full: how we count.

Editions

Every month, same method.

Each edition stays at its own URL after it is replaced, so a number you cited last month is still there to check. Only editions with complete or explicitly declared coverage are published.

September 2026 (this one) · August 2026 · July 2026

Method

How we count.

Plain GET requests (no rendering) against /.well-known/ai-catalog.json, /llms.txt, /agents.md and /robots.txt on each apex domain of the Tranco top-100K, paced at ~4 req/s globally, honoring robots.txt, as DesvelaBot/0.1 (bot policy). HTML responses to text paths count as absent. Every publisher is re-probed with a canary request to a nonexistent path: servers that answer 200 to anything are excluded as catch-alls, including some seemingly legitimate publishers. Undercounting beats inflating.

Full disclosure, and it cuts against us on purpose: our own domains (desvela.dev, desvela.ai) publish valid manifests, and they are deliberately not in any number on this page. They are not in the Tranco top-100K, so counting them would pad the very figures we are reporting. Same rule for every other domain we index outside the list. Verify ours with a GET like any other.
Subdomains are out of scope by design (docs.example.com publishing llms.txt does not count for example.com) — the census measures apex-domain adoption, same bar for everyone.

Read the full methodology → (universe, exclusions, coverage rules and what these numbers do not claim).

Stay on the list

Get next month's census in your inbox.

One email per month with the fresh numbers and what changed: new publishers, withdrawals, bot-blocking shifts. No product spam. Unsubscribe in one click from any email. We email you a confirmation link first, and nothing is sent until you click it.