The census is a measurement of adoption and access — which domains publish files an AI agent can read, and which bot directives they name or block. It is deliberately not a measurement of citations, recommendations or visibility inside AI systems. This page is the single reference for how every edition was collected.
Each edition crawls the top 100,000 domains of the Tranco list, taken on the day of the crawl. Eight surfaces per apex domain:
/.well-known/ard.json — the canonical ARD path since spec v0.91 (26 Aug 2026). Measured from 27 Aug 2026, one day after the spec moved./.well-known/ai-catalog.json — the predecessor ARD path, spec published May 2026. Demoted from MUST to MAY in v0.91 but still legal, and still the only one anyone publishes today. Measured since the first edition./llms.txt — the LLM-friendly site index./agents.md — agent instructions / capabilities file./robots.txt — the AI-bot directives (GPTBot, ClaudeBot, Google-Extended, PerplexityBot, …) named or blocked with Disallow: /./.well-known/agent-card.json — A2A agent card. Measured from 19 Aug 2026./.well-known/mcp.json — MCP server card. Measured from 19 Aug 2026.A surface is added, never substituted, and carries the date it entered. When the ARD spec moved its canonical path in August 2026, the obvious move was to swap one column for the other. We did not: a swap would have broken the month-over-month series at the exact point where the count of ARD publishers is single digits, and a reader comparing two editions would have seen a collapse that was ours, not the web's. Both paths are counted in parallel. A column that starts mid-series reads zero for the editions before it, and that zero means not measured, not nobody published — the dates above are how you tell the two apart.
Subdomains are out of scope by design: docs.example.com publishing an llms.txt does not count for example.com. The census measures apex-domain adoption, the same bar for everyone.
Plain GET requests, no rendering, against each path on each apex domain, paced at ~4 requests/second globally, honoring robots.txt, identified as DesvelaBot/0.1 (bot policy).
_index._agents.<domain> SVCB record, added to the spec 2026-07-02) is probed only on the monthly pass. A positive is discarded unless a wildcard canary for the same zone also fails.Only editions with complete or explicitly declared coverage are published. A crawl that does not finish is not silently reported as full: the generator refuses to emit a census unless at least 95% of the universe was crawled within the last 3 days, and a partial pass must declare its own denominator. An incomplete internal pass that never became a public edition does not appear in the archive, in downloads or in the editions list.
Each edition states, in its manifest, the number of domains checked, the universe size, when the data was collected, when it was published and when (if ever) it was corrected.
Each edition keeps its own URL, so a number cited last month is still there to check. Every published edition offers a downloadable JSON manifest (provenance + metrics) and a CSV of the tier aggregates.