Scraper.io

Connect your agent
28 sites··weekly·next run Mondays 04:40 UTC
CoveragePublish /llms.txtllmsPublish /llms-full.txtfullRefused by robots.txtrefusedTotal
6 of 13 category buckets · counts follow the searchclick a category to filter
SiteCategoryllms.txtSizellms-fullChangedSrc
28 of 28 sites · ordered by Tranco rank · each row opens the /llms.txt it was probed at

Do something with this data

Connect your agent

MCP-readable — wire it straight into Claude, Codex or Cursor.

npx -y mcp-remote https://fwgzmlpzreapyvvtfmpo.supabase.co/functions/v1/mcp --header "Authorization:Bearer $SCRAPERIO_API_KEY"Set up MCP

Get it as an API or CSV

Every column, every row, in the format your pipeline already reads.

Read the docs

Have an agent run on it

Spin up a long-running agent on this feed — apply, alert, enrich. Built and run by scraper.io.

Ask for a managed agent
Probed HEAD then GET under scraper.io feed-factory, one request per second per host, and only where the site’s robots.txt allows it · this feed shares ai-crawler-access-index’s crawl and made no network requests of its own on this run40-row slice of 400 probed · 6 publish /llms.txt and 4 publish /llms-full.txt in this slice · 3 were refused by robots.txt and are not counted either way · 5 carry a real Last-Modified; the rest date from the first run that saw those bytes · the category is derived from the domain, not stated by the site · feed id llms-txt-index12 withheld — failed validation (the origin answered nothing at all — DNS, TLS or connection failure)Agents: this feed is MCP-readable — connect at scraper.io/mcp

All rows, plain table

llms.txt Adoption Index: 400 records · updated 4 September 2026 · open — every row published here and in the CSV. which of the Tranco top 1,000 publish /llms.txt and /llms-full.txt, how big, and when they last changed

Last updated · run · 50 of 400 rows shown

llms.txt Adoption Index50 of 400 site records as of 4 September 2026. Every row links to the source it was read from.
#SiteCategoryllms.txtSize (KB)llms-fullChangedLast changedSource
1google.comsearchnonogoogle.com
2cloudflare.cominfrastructureyes16.8yes2026-09-04cloudflare.com
3gstatic.cominfrastructurerobots-refusedrobots-refusedgstatic.com
4facebook.comsocialrobots-refusedrobots-refusedfacebook.com
5microsoft.comothernono2025-09-18microsoft.com
6googleapis.cominfrastructurenonogoogleapis.com
7amazonaws.cominfrastructurenonoamazonaws.com
8youtube.commedianonoyoutube.com
9akamai.netinfrastructureunreachableunreachableakamai.net
10apple.comothernonoapple.com
11instagram.comsocialrobots-refusedrobots-refusedinstagram.com
12mail.ruotherrobots-refusedrobots-refusedmail.ru
13ezviz7.comotherunreachableunreachableezviz7.com
14fbcdn.netinfrastructurerobots-refusedrobots-refusedfbcdn.net
15dzen.ruothernonodzen.ru
16twitter.comsocialrobots-refusedrobots-refusedtwitter.com
17linkedin.comsocialnonolinkedin.com
18gtld-servers.netotherunreachableunreachablegtld-servers.net
19domaincontrol.comotherunreachableunreachabledomaincontrol.com
20googlevideo.comotherunreachableunreachablegooglevideo.com
21office.comothernonooffice.com
22googletagmanager.comadtechnonogoogletagmanager.com
23hicloudcam.comotherunreachableunreachablehicloudcam.com
24live.comothernonolive.com
25akamaiedge.netotherunreachableunreachableakamaiedge.net
26amazon.comcommercenonoamazon.com
27akadns.netotherunreachableunreachableakadns.net
28azure.comotheryes45.9no2026-02-17azure.com
29bing.comsearchnonobing.com
30github.comdeveloperyes28yes2026-09-04github.com
31wikipedia.orgreferencenonowikipedia.org
32whatsapp.netothernonowhatsapp.net
33apple-dns.netotherunreachableunreachableapple-dns.net
34googleusercontent.comothernonogoogleusercontent.com
35fastly.netinfrastructureyes15.4no2026-09-04fastly.net
36appsflyersdk.comotherrobots-refusedrobots-refusedappsflyersdk.com
37doubleclick.netadtechnonodoubleclick.net
38aaplimg.cominfrastructureunreachableunreachableaaplimg.com
39microsoftonline.comotherunreachableunreachablemicrosoftonline.com
40office.netotherunreachableunreachableoffice.net
41netflix.commediarobots-refusedrobots-refusednetflix.com
42trafficmanager.netotherunreachableunreachabletrafficmanager.net
43gandi.netothernonogandi.net
44sharepoint.comotherunreachableunreachablesharepoint.com
45youtu.beothernonoyoutu.be
46digicert.comotheryes8.2no2026-02-19digicert.com
47cloud.microsoftothernonocloud.microsoft
48wordpress.orgotheryes5.6no2026-09-04wordpress.org
49skype.comothernonoskype.com
50x.comsocialrobots-refusedrobots-refusedx.com

350 further records are in the downloads. Take the whole feed as CSV or JSON — open data, published under CC BY 4.0. Published by Scraper.io. These are the run’s rows as exported; the interactive view above applies each feed’s own quality filters on top of them.

llms.txt Adoption Index — Scraper.io