Scraper.io

Connect your agent
30 sites··weekly·next run Mondays 04:20 UTC
CoverageGPTBot blockedGPTClaudeBot blockedClaudePerplexity blockedPplxCCBot blockedCCTotal
4 of 4 stance buckets · counts follow the searchclick a stance to filter
SiteGPTBotClaudeBotPerplexityCCBotChangedSrc
30 of 30 sites · ordered by Tranco rank · each row opens the robots.txt it was read from

Do something with this data

Connect your agent

MCP-readable — wire it straight into Claude, Codex or Cursor.

npx -y mcp-remote https://fwgzmlpzreapyvvtfmpo.supabase.co/functions/v1/mcp --header "Authorization:Bearer $SCRAPERIO_API_KEY"Set up MCP

Get it as an API or CSV

Every column, every row, in the format your pipeline already reads.

Read the docs

Have an agent run on it

Spin up a long-running agent on this feed — apply, alert, enrich. Built and run by scraper.io.

Ask for a managed agent
Every verdict is read from the site’s own /robots.txt, fetched under the honest agent scraper.io feed-factory at one request per second per host · the directive line behind each verdict is published with it40-row slice of 400 in this run · 12 name at least one AI crawler in their own rule group · 3 block only through the wildcard group · 9 publish no robots.txt at all · first run is a baseline, so every row reads baseline rather than claiming a change · feed id ai-crawler-access-index10 withheld — failed validation (the origin answered nothing at all — DNS, TLS or connection failure)Agents: this feed is MCP-readable — connect at scraper.io/mcp

All rows, plain table

AI Crawler Access Index: 400 records · updated baseline · open — every row published here and in the CSV. robots.txt verdicts for nine AI crawlers across the Tranco top 1,000, with the directive line behind each

Last updated · run · 50 of 400 rows shown

AI Crawler Access Index50 of 400 site records as of baseline. Every row links to the source it was read from.
#SiteGPTBotClaudeBotPerplexityCCBotChangedLast changedSource
1google.comallowedallowedallowedallowedbaselinegoogle.com
2cloudflare.comallowedallowedallowedallowedbaselinecloudflare.com
3gstatic.comblockedblockedblockedblockedbaselinegstatic.com
4facebook.comallowedallowedallowedblockedbaselinefacebook.com
5microsoft.comallowedallowedallowedallowedbaselinemicrosoft.com
6googleapis.comno robotsno robotsno robotsno robotsbaselinegoogleapis.com
7amazonaws.comallowedallowedallowedallowedbaselineamazonaws.com
8youtube.comallowedallowedallowedallowedbaselineyoutube.com
9akamai.netunreachableunreachableunreachableunreachablebaselineakamai.net
10apple.comallowedallowedallowedallowedbaselineapple.com
11instagram.comblockedblockedblockedblockedbaselineinstagram.com
12mail.ruallowedallowedallowedallowedbaselinemail.ru
13ezviz7.comunreachableunreachableunreachableunreachablebaselineezviz7.com
14fbcdn.netblockedblockedblockedblockedbaselinefbcdn.net
15dzen.ruallowedallowedallowedallowedbaselinedzen.ru
16twitter.comblockedblockedblockedblockedbaselinetwitter.com
17linkedin.comno robotsno robotsno robotsno robotsbaselinelinkedin.com
18gtld-servers.netunreachableunreachableunreachableunreachablebaselinegtld-servers.net
19domaincontrol.comunreachableunreachableunreachableunreachablebaselinedomaincontrol.com
20googlevideo.comunreachableunreachableunreachableunreachablebaselinegooglevideo.com
21office.comno robotsno robotsno robotsno robotsbaselineoffice.com
22googletagmanager.comno robotsno robotsno robotsno robotsbaselinegoogletagmanager.com
23hicloudcam.comunreachableunreachableunreachableunreachablebaselinehicloudcam.com
24live.comno robotsno robotsno robotsno robotsbaselinelive.com
25akamaiedge.netunreachableunreachableunreachableunreachablebaselineakamaiedge.net
26amazon.comblockedblockedblockedblockedbaselineamazon.com
27akadns.netunreachableunreachableunreachableunreachablebaselineakadns.net
28azure.comno robotsno robotsno robotsno robotsbaselineazure.com
29bing.comallowedallowedallowedallowedbaselinebing.com
30github.comallowedallowedallowedallowedbaselinegithub.com
31wikipedia.orgallowedallowedallowedallowedbaselinewikipedia.org
32whatsapp.netblockedblockedblockedallowedbaselinewhatsapp.net
33apple-dns.netunreachableunreachableunreachableunreachablebaselineapple-dns.net
34googleusercontent.comno robotsno robotsno robotsno robotsbaselinegoogleusercontent.com
35fastly.netallowedallowedallowedallowedbaselinefastly.net
36appsflyersdk.comblockedblockedblockedblockedbaselineappsflyersdk.com
37doubleclick.netallowedallowedallowedallowedbaselinedoubleclick.net
38aaplimg.comunreachableunreachableunreachableunreachablebaselineaaplimg.com
39microsoftonline.comunreachableunreachableunreachableunreachablebaselinemicrosoftonline.com
40office.netunreachableunreachableunreachableunreachablebaselineoffice.net
41netflix.comallowedallowedallowedblockedbaselinenetflix.com
42trafficmanager.netunreachableunreachableunreachableunreachablebaselinetrafficmanager.net
43gandi.netallowedallowedallowedallowedbaselinegandi.net
44sharepoint.comunreachableunreachableunreachableunreachablebaselinesharepoint.com
45youtu.beallowedallowedallowedallowedbaselineyoutu.be
46digicert.comallowedallowedallowedallowedbaselinedigicert.com
47cloud.microsoftno robotsno robotsno robotsno robotsbaselinecloud.microsoft
48wordpress.orgallowedallowedallowedallowedbaselinewordpress.org
49skype.comno robotsno robotsno robotsno robotsbaselineskype.com
50x.comblockedblockedblockedblockedbaselinex.com

350 further records are in the downloads. Take the whole feed as CSV or JSON — open data, published under CC BY 4.0. Published by Scraper.io. These are the run’s rows as exported; the interactive view above applies each feed’s own quality filters on top of them.