State of AI access
Once a week I read what the top 1,000 websites tell AI crawlers: who gets blocked, who publishes llms.txt, and who has worked out how to ask for money.
Data through 2026-09-30 · updates weekly
CCBot is the most blocked AI crawler: 19% of top sites with a robots.txt block it by name
Share of the Tranco top 998 domains with a robots.txt (482) whose rules name the crawler and disallow the whole site. 313 list entries are CDN, API or infrastructure hostnames with no reachable robots.txt; they are excluded.
Pricing is almost absent: 1 of 685 top sites state a machine-readable price; 15% publish llms.txt
Adoption of machine-readable access signals among top sites that serve a website. Block rows use sites with a robots.txt as the denominator.
Top sites vs the whole crawl
My earlier Common Crawl census (CC-MAIN-2026-39, a 5% sample of robots.txt files, 2.67M hosts) measured the same rules across the whole web. Method break: that population is every crawled host, this one is the top 1,000 registrable domains; compare levels with care, and read the weekly series from 2026-09-30 on as the consistent one.