RadiusRankResearch

Original RadiusRank research

Only 27 of 250 sites named an AI crawler in robots.txt.

We fetched the homepage, robots.txt and llms.txt of 250 public websites on August 14, 2026 — 50 each across SaaS, AI prosumer tools, creative agencies, ecommerce and DTC brands, and law firms — and ran identical deterministic checks over all of them. 11% named any AI crawler at all, and the strongest cohort (SaaS) still only reached 8 of 50.

Updated 250-site dataset across five cohorts

Key findings

  • 27 of 250 (11%) named any AI crawler in robots.txt. This is the one finding that holds across every cohort rather than separating them.
  • llms.txt is where cohorts diverge most, from 6 of 50 (Law firm) to 31 of 50.
  • Two cohorts exposed no product or service schema at all. Creative agencies and law firms both returned zero, so nothing on those homepages described what the business sells in a machine-readable way.

All five cohorts

Every cohort was sampled at n=50 and parsed with the same rules, so the columns are directly comparable. Figures are out of 50.

SignalSaaSAI prosumer toolsCreative agenciesEcommerce and DTCLaw firms
Canonical URL exposed3128323839
Organization schema present2719173024
Product or service schema present1611020
Named an AI crawler in robots.txt88623
Served a valid llms.txt31259286

SaaS

Of 50 saas sampled, 48 returned a homepage to a research crawler. 31 exposed a canonical URL, 27 carried Organization schema, 16 carried product or service schema, 8 named an AI crawler in robots.txt, and 31 served a valid llms.txt.

2 refused the research user-agent (formbricks.com, gusto.com) and are counted as absent rather than dropped from the denominator.

AI prosumer tools

Of 50 ai prosumer tools sampled, 40 returned a homepage to a research crawler. 28 exposed a canonical URL, 19 carried Organization schema, 11 carried product or service schema, 8 named an AI crawler in robots.txt, and 25 served a valid llms.txt.

10 refused the research user-agent (gamma.app, ideogram.ai, leonardo.ai, lovable.dev, make.com, openai.com, canva.com, capcut.com, midjourney.com, perplexity.ai) and are counted as absent rather than dropped from the denominator.

Creative agencies

Of 50 creative agencies sampled, 48 returned a homepage to a research crawler. 32 exposed a canonical URL, 17 carried Organization schema, 0 carried product or service schema, 6 named an AI crawler in robots.txt, and 9 served a valid llms.txt.

2 refused the research user-agent (thecreativeagencyco.com, pearlfisher.com) and are counted as absent rather than dropped from the denominator.

Ecommerce and DTC

Of 50 ecommerce and dtc sampled, 42 returned a homepage to a research crawler. 38 exposed a canonical URL, 30 carried Organization schema, 2 carried product or service schema, 2 named an AI crawler in robots.txt, and 28 served a valid llms.txt.

8 refused the research user-agent (athleticgreens.com, bombas.com, gymshark.com, hims.com, hydroflask.com, patagonia.com, ro.co, savagex.com) and are counted as absent rather than dropped from the denominator.

Law firms

Of 50 law firms sampled, 44 returned a homepage to a research crawler. 39 exposed a canonical URL, 24 carried Organization schema, 0 carried product or service schema, 3 named an AI crawler in robots.txt, and 6 served a valid llms.txt.

6 refused the research user-agent (akingump.com, dlapiper.com, jenner.com, orrick.com, ropesgray.com, sheppardmullin.com) and are counted as absent rather than dropped from the denominator.

Methodology

The SaaS cohort is Y Combinator-backed and Product Hunt-launched B2B SaaS, not the enterprise brands an earlier run used, because the buyers this study is for look like the former. The law firm cohort is 50 law firms only; an earlier run mixed law, accounting and consulting under one label, which meant the figure could not be attributed to any single industry. Each site was requested once over HTTPS with a declared research user-agent, following redirects. The homepage, /robots.txt and /llms.txt were fetched and parsed with deterministic rules. No paid SEO provider, LLM, analytics or Search Console data was used. Every host was probed beforehand to confirm it answers; hosts that refuse a research user-agent remain in the sample and are recorded as failed fetches.

Requests used the user-agent RadiusRankResearchBot/1.0 (+https://www.radiusrank.com/research; research@radiusrank.com), followed redirects, and timed out after 12 seconds. Each site received exactly three requests. No paid SEO provider, LLM, analytics or Search Console data was used, and the generator is committed at scripts/generate-cohort-readiness-benchmark.mjs.

Limitations

  • Homepage only. A site may expose schema on interior pages this study did not fetch.
  • 28 of 250 sites refused the research user-agent. They stay in the denominator, so their signals count as absent rather than being dropped.
  • Each cohort is a named, non-random sample of recognisable companies, not a random draw from its industry.
  • The study measures whether a signal is present, not whether an answer engine acted on it. It cannot show causation with citations or traffic.
  • RadiusRank sells AI SEO services and has a commercial interest in the category this study measures.

Data

The full record set for all 250 sites, including per-site HTTP status and homepage hashes, is published as cohort-readiness-2026-08-14.json.

Start with the evidence

Run the same checks on your own site.

Every signal in this study is deterministic and public. You can verify each one on your own domain before deciding what to fix.

Run the free diagnostic