RadiusRankResearch

Original RadiusRank research

Only 23% of readable robots.txt files named an AI agent.

We inspected public robots.txt files for 50 AI, SEO, local-search, and CMS brands on August 11, 2026. 47 files were readable. Only 11 of those 47 files explicitly named any of 11 AI-related agents. Most sites left AI access to general wildcard rules rather than making a crawler-specific choice.

Updated 50-site dataset and methodology

Short answerCrawler-specific control is still uncommon in this selected technology sample. Explicit configuration appeared on 11 of 47 readable robots files. Retrieval agents were not blocked at the root on any readable site, while a small number of sites selectively blocked training or AI-control tokens.
50selected technology sites sampled
47readable robots.txt files
23%named at least one AI-related agent
87%declared a sitemap in robots.txt

What the 50-site snapshot found

  1. Explicit AI controls were the exception. Eleven of 47 readable files named at least one studied AI agent. Thirty-six readable files relied on wildcard rules or had no matching rules for those agents.
  2. GPTBot was the most frequently named training crawler. Eight readable files named GPTBot. Seven named OAI-SearchBot, showing that many teams that configure OpenAI controls distinguish training from search retrieval.
  3. Search and user-directed access stayed open at the root. None of the readable files blocked OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, or Perplexity-User at the root path under the rules we parsed.
  4. Selective blocks concentrated on training and control tokens.GPTBot, ClaudeBot, CCBot, and Applebot-Extended were each blocked at the root by one readable file. Google-Extended was blocked by two.
  5. Three domains could not be classified. Similarweb failed during collection, while OtterlyAI and AEO Engine returned HTTP 403. We report those checks as unavailable instead of assuming access.
This is a descriptive benchmark of a selected set of 50 relevant technology brands, including RadiusRank. It is not a random sample and should not be interpreted as a market-wide prevalence estimate.

Training, search, and user-directed agents are different controls

A single AI-crawler score hides the most important distinction. OpenAI, Anthropic, and Perplexity publish separate agents for different uses. A site may permit search discovery while declining model-training access. Site owners should make those choices independently.

PurposeAgents in this studyWhat the rule controls
Training or dataset collectionGPTBot, ClaudeBot, CCBotCollection that can contribute to model or public training corpora
Search retrievalOAI-SearchBot, Claude-SearchBot, PerplexityBotDiscovery or indexing used to support answers and search results
User-directed retrievalChatGPT-User, Claude-User, Perplexity-UserA page visit initiated in response to a user request
Product-level AI controlsGoogle-Extended, Applebot-ExtendedHow already-crawled content may be used by specified AI products

Perplexity states that Perplexity-User generally ignores robots.txt because the fetch is user requested. Its row is still useful for detecting whether publishers name the token, but the rule should not be treated as proof of enforcement.

The 50-site AI crawler access dataset

“Allowed” below means the applicable robots.txt group did not block the domain root, /, for that token. It does not prove that every path is crawlable, that the crawler will visit the site, or that content will be indexed, trained on, or cited.

Allowed at root Blocked at root Check unavailable
BrandExplicit AI agentsGPTBotOAI SearchClaudeBotClaude SearchPerplexityBotGoogle Extended
OpenAIopenai.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Anthropicanthropic.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Perplexityperplexity.aiNone explicitly namedallowedallowedallowedallowedallowedallowed
Coherecohere.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Mistral AImistral.aiNone explicitly namedallowedallowedallowedallowedallowedallowed
Hugging Facehuggingface.coNone explicitly namedallowedallowedallowedallowedallowedallowed
Together AItogether.aiGoogle-Extendedallowedallowedallowedallowedallowedblocked
Replicatereplicate.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Groqgroq.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Ahrefsahrefs.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Semrushsemrush.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Similarwebsimilarweb.comNone explicitly namedunavailableunavailableunavailableunavailableunavailableunavailable
SE Rankingseranking.comGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extendedallowedallowedallowedallowedallowedallowed
Mozmoz.comGPTBotallowedallowedallowedallowedallowedallowed
SISTRIXsistrix.comNone explicitly namedallowedallowedallowedallowedallowedallowed
BrightEdgebrightedge.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Majesticmajestic.comNone explicitly namedallowedallowedallowedallowedallowedallowed
SpyFuspyfu.comChatGPT-User, Claude-User, Perplexity-Userallowedallowedallowedallowedallowedallowed
Seobilityseobility.netGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Applebot-Extendedallowedallowedallowedallowedallowedallowed
Mangoolsmangools.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Screaming Frogscreamingfrog.co.ukNone explicitly namedallowedallowedallowedallowedallowedallowed
Sitebulbsitebulb.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Botifybotify.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Oncrawloncrawl.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Surfersurferseo.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Clearscopeclearscope.ioNone explicitly namedallowedallowedallowedallowedallowedallowed
MarketMusemarketmuse.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Frasefrase.ioGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, CCBot, Applebot-Extendedallowedallowedallowedallowedallowedallowed
Jasperjasper.aiGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extendedallowedallowedallowedallowedallowedallowed
Copy.aicopy.aiNone explicitly namedallowedallowedallowedallowedallowedallowed
Writesonicwritesonic.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Scalenutscalenut.comNone explicitly namedallowedallowedallowedallowedallowedallowed
NeuronWriterneuronwriter.comNone explicitly namedallowedallowedallowedallowedallowedallowed
SEOmaticseomatic.aiOAI-SearchBot, Claude-SearchBot, PerplexityBotallowedallowedallowedallowedallowedallowed
Profoundtryprofound.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Peec AIpeec.aiNone explicitly namedallowedallowedallowedallowedallowedallowed
OtterlyAIotterly.aiNone explicitly namedunavailableunavailableunavailableunavailableunavailableunavailable
Scrunch AIscrunchai.comNone explicitly namedallowedallowedallowedallowedallowedallowed
AEO Engineaeoengine.aiNone explicitly namedunavailableunavailableunavailableunavailableunavailableunavailable
RadiusRankradiusrank.comGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, CCBot, Applebot-Extendedallowedallowedallowedallowedallowedallowed
Yextyext.comGPTBot, ClaudeBot, PerplexityBot, CCBotallowedallowedallowedallowedallowedallowed
BrightLocalbrightlocal.comGPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended, CCBot, Applebot-Extendedblockedallowedblockedallowedallowedblocked
Whitesparkwhitespark.caNone explicitly namedallowedallowedallowedallowedallowedallowed
Local Falconlocalfalcon.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Local Vikinglocalviking.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Yoastyoast.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Rank Mathrankmath.comNone explicitly namedallowedallowedallowedallowedallowedallowed
All in One SEOaioseo.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Webflowwebflow.comNone explicitly namedallowedallowedallowedallowedallowedallowed
Wixwix.comNone explicitly namedallowedallowedallowedallowedallowedallowed

Methodology and reproducibility

The sample contains 50 known companies across AI platforms, SEO platforms, content optimization, AI visibility, local SEO, and CMS software. The list was selected for relevance to the search and AI-discovery market. It was not randomly drawn.

  1. Request https://domain/robots.txt with a named RadiusRank research user agent and a 15-second timeout.
  2. Follow redirects and record the final URL, HTTP status, content type, and byte size.
  3. Reject HTTP errors, network failures, and HTML responses rather than interpreting them as valid robots files.
  4. Parse user-agent groups, Allow, Disallow, wildcards, end anchors, and Sitemap declarations.
  5. Choose the most specific matching agent group, then apply the longest matching path rule at /, with Allow winning exact ties.
  6. Store the SHA-256 hash of every readable source so future snapshots can detect changes without republishing third-party files.
BENCHMARK_INPUT=ai-crawler-benchmark-domains.csv BENCHMARK_DATE=2026-08-11 node ai-crawler-access-benchmark-generator.mjs

The downloadable generator uses Node.js built-ins and the source domain list in this repository. Run it again with a new date to create a fresh CSV and JSON snapshot.

What this benchmark does not prove

  • Root access is not whole-site access. A site may allow / while blocking private paths, search pages, parameters, media, or specific directories.
  • robots.txt is a declared preference. The benchmark does not verify crawler IP addresses, actual requests, compliance, indexing, model training, citations, or rankings.
  • A 403 is not classified as blocked by robots.txt. It may come from a firewall, bot defense, authentication, or another server policy.
  • The web changes. Results describe the collected snapshot. Source hashes and the collection date are included so changes can be audited.
  • The sample is intentionally narrow. These results describe selected technology brands and cannot estimate behavior across the whole web.

What site owners should do

  1. Test the effective rules. Use the free robots.txt checker on production content paths, not only the home page.
  2. Separate policy decisions. Decide independently whether to allow training, search retrieval, product controls, and user-requested visits.
  3. Generate narrowly scoped rules. The robots.txt generator includes separate AI-agent controls.
  4. Check server enforcement. Review CDN and WAF logs because a correct robots file cannot fix a firewall that blocks the real crawler.
  5. Monitor changes. Store dated copies or hashes of robots.txt so accidental policy changes are visible.

Crawler access is only one part of answer engine optimization. Read the practical AEO guide or see how RadiusRank handles technical eligibility, evidence, and authority in its answer engine optimization service.

Primary documentation

You may cite the benchmark with the snapshot date and a link to this page. For a custom breakdown or a correction supported by a public source, contact team@radiusrank.com.

Start with the evidence

Check your own AI crawler rules.

Use the same parser on your domain, then decide separately which training, search, and user-directed agents should have access.

Run the free diagnostic