Key findings
- 27 of 250 (11%) named any AI crawler in robots.txt. This is the one finding that holds across every cohort rather than separating them.
- llms.txt is where cohorts diverge most, from 6 of 50 (Law firm) to 31 of 50.
- Two cohorts exposed no product or service schema at all. Creative agencies and law firms both returned zero, so nothing on those homepages described what the business sells in a machine-readable way.
All five cohorts
Every cohort was sampled at n=50 and parsed with the same rules, so the columns are directly comparable. Figures are out of 50.
| Signal | SaaS | AI prosumer tools | Creative agencies | Ecommerce and DTC | Law firms |
|---|---|---|---|---|---|
| Canonical URL exposed | 31 | 28 | 32 | 38 | 39 |
| Organization schema present | 27 | 19 | 17 | 30 | 24 |
| Product or service schema present | 16 | 11 | 0 | 2 | 0 |
| Named an AI crawler in robots.txt | 8 | 8 | 6 | 2 | 3 |
| Served a valid llms.txt | 31 | 25 | 9 | 28 | 6 |
SaaS
Of 50 saas sampled, 48 returned a homepage to a research crawler. 31 exposed a canonical URL, 27 carried Organization schema, 16 carried product or service schema, 8 named an AI crawler in robots.txt, and 31 served a valid llms.txt.
2 refused the research user-agent (formbricks.com, gusto.com) and are counted as absent rather than dropped from the denominator.
AI prosumer tools
Of 50 ai prosumer tools sampled, 40 returned a homepage to a research crawler. 28 exposed a canonical URL, 19 carried Organization schema, 11 carried product or service schema, 8 named an AI crawler in robots.txt, and 25 served a valid llms.txt.
10 refused the research user-agent (gamma.app, ideogram.ai, leonardo.ai, lovable.dev, make.com, openai.com, canva.com, capcut.com, midjourney.com, perplexity.ai) and are counted as absent rather than dropped from the denominator.
Creative agencies
Of 50 creative agencies sampled, 48 returned a homepage to a research crawler. 32 exposed a canonical URL, 17 carried Organization schema, 0 carried product or service schema, 6 named an AI crawler in robots.txt, and 9 served a valid llms.txt.
2 refused the research user-agent (thecreativeagencyco.com, pearlfisher.com) and are counted as absent rather than dropped from the denominator.
Ecommerce and DTC
Of 50 ecommerce and dtc sampled, 42 returned a homepage to a research crawler. 38 exposed a canonical URL, 30 carried Organization schema, 2 carried product or service schema, 2 named an AI crawler in robots.txt, and 28 served a valid llms.txt.
8 refused the research user-agent (athleticgreens.com, bombas.com, gymshark.com, hims.com, hydroflask.com, patagonia.com, ro.co, savagex.com) and are counted as absent rather than dropped from the denominator.
Law firms
Of 50 law firms sampled, 44 returned a homepage to a research crawler. 39 exposed a canonical URL, 24 carried Organization schema, 0 carried product or service schema, 3 named an AI crawler in robots.txt, and 6 served a valid llms.txt.
6 refused the research user-agent (akingump.com, dlapiper.com, jenner.com, orrick.com, ropesgray.com, sheppardmullin.com) and are counted as absent rather than dropped from the denominator.
Methodology
The SaaS cohort is Y Combinator-backed and Product Hunt-launched B2B SaaS, not the enterprise brands an earlier run used, because the buyers this study is for look like the former. The law firm cohort is 50 law firms only; an earlier run mixed law, accounting and consulting under one label, which meant the figure could not be attributed to any single industry. Each site was requested once over HTTPS with a declared research user-agent, following redirects. The homepage, /robots.txt and /llms.txt were fetched and parsed with deterministic rules. No paid SEO provider, LLM, analytics or Search Console data was used. Every host was probed beforehand to confirm it answers; hosts that refuse a research user-agent remain in the sample and are recorded as failed fetches.
Requests used the user-agent RadiusRankResearchBot/1.0 (+https://www.radiusrank.com/research; research@radiusrank.com), followed redirects, and timed out after 12 seconds. Each site received exactly three requests. No paid SEO provider, LLM, analytics or Search Console data was used, and the generator is committed at scripts/generate-cohort-readiness-benchmark.mjs.
Limitations
- Homepage only. A site may expose schema on interior pages this study did not fetch.
- 28 of 250 sites refused the research user-agent. They stay in the denominator, so their signals count as absent rather than being dropped.
- Each cohort is a named, non-random sample of recognisable companies, not a random draw from its industry.
- The study measures whether a signal is present, not whether an answer engine acted on it. It cannot show causation with citations or traffic.
- RadiusRank sells AI SEO services and has a commercial interest in the category this study measures.
Data
The full record set for all 250 sites, including per-site HTTP status and homepage hashes, is published as cohort-readiness-2026-08-14.json.