What the 50-site snapshot found
- Explicit AI controls were the exception. Eleven of 47 readable files named at least one studied AI agent. Thirty-six readable files relied on wildcard rules or had no matching rules for those agents.
- GPTBot was the most frequently named training crawler. Eight readable files named GPTBot. Seven named OAI-SearchBot, showing that many teams that configure OpenAI controls distinguish training from search retrieval.
- Search and user-directed access stayed open at the root. None of the readable files blocked OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, or Perplexity-User at the root path under the rules we parsed.
- Selective blocks concentrated on training and control tokens.GPTBot, ClaudeBot, CCBot, and Applebot-Extended were each blocked at the root by one readable file. Google-Extended was blocked by two.
- Three domains could not be classified. Similarweb failed during collection, while OtterlyAI and AEO Engine returned HTTP 403. We report those checks as unavailable instead of assuming access.
Training, search, and user-directed agents are different controls
A single AI-crawler score hides the most important distinction. OpenAI, Anthropic, and Perplexity publish separate agents for different uses. A site may permit search discovery while declining model-training access. Site owners should make those choices independently.
| Purpose | Agents in this study | What the rule controls |
|---|---|---|
| Training or dataset collection | GPTBot, ClaudeBot, CCBot | Collection that can contribute to model or public training corpora |
| Search retrieval | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Discovery or indexing used to support answers and search results |
| User-directed retrieval | ChatGPT-User, Claude-User, Perplexity-User | A page visit initiated in response to a user request |
| Product-level AI controls | Google-Extended, Applebot-Extended | How already-crawled content may be used by specified AI products |
Perplexity states that Perplexity-User generally ignores robots.txt because the fetch is user requested. Its row is still useful for detecting whether publishers name the token, but the rule should not be treated as proof of enforcement.
The 50-site AI crawler access dataset
“Allowed” below means the applicable robots.txt group did not block the domain root, /, for that token. It does not prove that every path is crawlable, that the crawler will visit the site, or that content will be indexed, trained on, or cited.
| Brand | Explicit AI agents | GPTBot | OAI Search | ClaudeBot | Claude Search | PerplexityBot | Google Extended |
|---|---|---|---|---|---|---|---|
| OpenAIopenai.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Anthropicanthropic.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Perplexityperplexity.ai | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Coherecohere.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Mistral AImistral.ai | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Hugging Facehuggingface.co | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Together AItogether.ai | Google-Extended | allowed | allowed | allowed | allowed | allowed | blocked |
| Replicatereplicate.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Groqgroq.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Ahrefsahrefs.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Semrushsemrush.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Similarwebsimilarweb.com | None explicitly named | unavailable | unavailable | unavailable | unavailable | unavailable | unavailable |
| SE Rankingseranking.com | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended | allowed | allowed | allowed | allowed | allowed | allowed |
| Mozmoz.com | GPTBot | allowed | allowed | allowed | allowed | allowed | allowed |
| SISTRIXsistrix.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| BrightEdgebrightedge.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Majesticmajestic.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| SpyFuspyfu.com | ChatGPT-User, Claude-User, Perplexity-User | allowed | allowed | allowed | allowed | allowed | allowed |
| Seobilityseobility.net | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Applebot-Extended | allowed | allowed | allowed | allowed | allowed | allowed |
| Mangoolsmangools.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Screaming Frogscreamingfrog.co.uk | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Sitebulbsitebulb.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Botifybotify.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Oncrawloncrawl.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Surfersurferseo.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Clearscopeclearscope.io | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| MarketMusemarketmuse.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Frasefrase.io | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, CCBot, Applebot-Extended | allowed | allowed | allowed | allowed | allowed | allowed |
| Jasperjasper.ai | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, CCBot, Applebot-Extended | allowed | allowed | allowed | allowed | allowed | allowed |
| Copy.aicopy.ai | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Writesonicwritesonic.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Scalenutscalenut.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| NeuronWriterneuronwriter.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| SEOmaticseomatic.ai | OAI-SearchBot, Claude-SearchBot, PerplexityBot | allowed | allowed | allowed | allowed | allowed | allowed |
| Profoundtryprofound.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Peec AIpeec.ai | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| OtterlyAIotterly.ai | None explicitly named | unavailable | unavailable | unavailable | unavailable | unavailable | unavailable |
| Scrunch AIscrunchai.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| AEO Engineaeoengine.ai | None explicitly named | unavailable | unavailable | unavailable | unavailable | unavailable | unavailable |
| RadiusRankradiusrank.com | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, CCBot, Applebot-Extended | allowed | allowed | allowed | allowed | allowed | allowed |
| Yextyext.com | GPTBot, ClaudeBot, PerplexityBot, CCBot | allowed | allowed | allowed | allowed | allowed | allowed |
| BrightLocalbrightlocal.com | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended, CCBot, Applebot-Extended | blocked | allowed | blocked | allowed | allowed | blocked |
| Whitesparkwhitespark.ca | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Local Falconlocalfalcon.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Local Vikinglocalviking.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Yoastyoast.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Rank Mathrankmath.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| All in One SEOaioseo.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Webflowwebflow.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
| Wixwix.com | None explicitly named | allowed | allowed | allowed | allowed | allowed | allowed |
Methodology and reproducibility
The sample contains 50 known companies across AI platforms, SEO platforms, content optimization, AI visibility, local SEO, and CMS software. The list was selected for relevance to the search and AI-discovery market. It was not randomly drawn.
- Request
https://domain/robots.txtwith a named RadiusRank research user agent and a 15-second timeout. - Follow redirects and record the final URL, HTTP status, content type, and byte size.
- Reject HTTP errors, network failures, and HTML responses rather than interpreting them as valid robots files.
- Parse user-agent groups, Allow, Disallow, wildcards, end anchors, and Sitemap declarations.
- Choose the most specific matching agent group, then apply the longest matching path rule at
/, with Allow winning exact ties. - Store the SHA-256 hash of every readable source so future snapshots can detect changes without republishing third-party files.
BENCHMARK_INPUT=ai-crawler-benchmark-domains.csv BENCHMARK_DATE=2026-08-11 node ai-crawler-access-benchmark-generator.mjsThe downloadable generator uses Node.js built-ins and the source domain list in this repository. Run it again with a new date to create a fresh CSV and JSON snapshot.
What this benchmark does not prove
- Root access is not whole-site access. A site may allow
/while blocking private paths, search pages, parameters, media, or specific directories. - robots.txt is a declared preference. The benchmark does not verify crawler IP addresses, actual requests, compliance, indexing, model training, citations, or rankings.
- A 403 is not classified as blocked by robots.txt. It may come from a firewall, bot defense, authentication, or another server policy.
- The web changes. Results describe the collected snapshot. Source hashes and the collection date are included so changes can be audited.
- The sample is intentionally narrow. These results describe selected technology brands and cannot estimate behavior across the whole web.
What site owners should do
- Test the effective rules. Use the free robots.txt checker on production content paths, not only the home page.
- Separate policy decisions. Decide independently whether to allow training, search retrieval, product controls, and user-requested visits.
- Generate narrowly scoped rules. The robots.txt generator includes separate AI-agent controls.
- Check server enforcement. Review CDN and WAF logs because a correct robots file cannot fix a firewall that blocks the real crawler.
- Monitor changes. Store dated copies or hashes of robots.txt so accidental policy changes are visible.
Crawler access is only one part of answer engine optimization. Read the practical AEO guide or see how RadiusRank handles technical eligibility, evidence, and authority in its answer engine optimization service.
Primary documentation
- OpenAI crawler documentation: GPTBot, OAI-SearchBot, and ChatGPT-User
- Anthropic crawler documentation: ClaudeBot, Claude-SearchBot, and Claude-User
- Perplexity crawler documentation: PerplexityBot and Perplexity-User
- Google crawler documentation: Google-Extended
- Google Search Central: introduction to robots.txt
You may cite the benchmark with the snapshot date and a link to this page. For a custom breakdown or a correction supported by a public source, contact team@radiusrank.com.