Key findings
- The gap is not effort, it is vintage. 91 of 100 sites point robots.txt at a sitemap; only 15 name a single AI crawler. These teams maintain their robots file. They have simply never added the agents that read it now.
- 23 of 100 (23%) declare themselves as software. Every company in this sample sells a product or an app, but only this many say so in a machine-readable way, even though 60 manage Organization schema.
- Only 3 of 100 carried all five signals — inngest.com, render.com, weaviate.io.
- llms.txt is no longer the rare one. 67 of 100 served a valid file, which is more than managed Organization schema. The convention arrived faster in this cohort than the older ones did.
Signal by signal
| Signal | Sites | Share |
|---|---|---|
| Homepage returned to a research crawler | 96 / 100 | 96% |
| Title tag | 98 / 100 | 98% |
| Meta description | 77 / 100 | 77% |
| robots.txt points at a sitemap | 91 / 100 | 91% |
| Any JSON-LD at all | 66 / 100 | 66% |
| Canonical URL | 68 / 100 | 68% |
| Organization schema | 60 / 100 | 60% |
| Product or app schema | 23 / 100 | 23% |
| Named an AI crawler | 15 / 100 | 15% |
| Valid llms.txt | 67 / 100 | 67% |
| FAQ schema | 6 / 100 | 6% |
How many signals per site
An average would hide the shape. Scored against the five signals above, the sample clusters low: most sites carry two.
- 0 of 59
- 1 of 513
- 2 of 536
- 3 of 523
- 4 of 516
- 5 of 53
Which AI agents get named
Across the 15 sites that named any agent, these are the ones that appear. A site naming an agent has made a deliberate choice about it, whether that choice is allow or disallow.
| Agent | Sites naming it |
|---|---|
PerplexityBot | 12 |
GPTBot | 11 |
ClaudeBot | 11 |
ChatGPT-User | 10 |
CCBot | 10 |
OAI-SearchBot | 8 |
Google-Extended | 8 |
Bytespider | 6 |
Applebot-Extended | 5 |
Perplexity-User | 3 |
meta-externalagent | 3 |
Claude-SearchBot | 3 |
Claude-User | 2 |
All 100 sites
Ordered by how many of the five signals each carries. Every company in the sample is named, because a benchmark you cannot audit is not evidence.
| Company | Canonical URL | Organization schema | Product or app schema | Named an AI crawler | Valid llms.txt | Score |
|---|---|---|---|---|---|---|
| inngest.com | yes | yes | yes | yes | yes | 5 |
| render.com | yes | yes | yes | yes | yes | 5 |
| weaviate.io | yes | yes | yes | yes | yes | 5 |
| airbyte.com | yes | yes | yes | no | yes | 4 |
| amplitude.com | yes | yes | no | yes | yes | 4 |
| atlan.com | yes | yes | no | yes | yes | 4 |
| attio.com | no | yes | yes | yes | yes | 4 |
| calendly.com | yes | yes | no | yes | yes | 4 |
| clerk.com | yes | yes | no | yes | yes | 4 |
| folk.app | no | yes | yes | yes | yes | 4 |
| guidde.com | yes | yes | yes | yes | no | 4 |
| knock.app | yes | yes | yes | no | yes | 4 |
| novu.co | yes | yes | yes | no | yes | 4 |
| railway.app | yes | yes | yes | no | yes | 4 |
| resend.com | yes | yes | yes | no | yes | 4 |
| rootly.com | no | yes | yes | yes | yes | 4 |
| tally.so | yes | yes | yes | no | yes | 4 |
| temporal.io | yes | yes | yes | no | yes | 4 |
| vercel.com | yes | yes | yes | no | yes | 4 |
| arcade.software | no | yes | yes | no | yes | 3 |
| astronomer.io | yes | yes | no | no | yes | 3 |
| brex.com | yes | yes | no | no | yes | 3 |
| budibase.com | yes | yes | no | no | yes | 3 |
| chronosphere.io | yes | yes | no | no | yes | 3 |
| comet.com | yes | yes | no | no | yes | 3 |
| deel.com | yes | no | no | yes | yes | 3 |
| figma.com | yes | yes | no | yes | no | 3 |
| framer.com | yes | yes | no | no | yes | 3 |
| front.com | yes | yes | no | no | yes | 3 |
| fullstory.com | yes | yes | no | no | yes | 3 |
| grafana.com | yes | yes | no | no | yes | 3 |
| hex.tech | yes | yes | yes | no | no | 3 |
| incident.io | yes | yes | no | no | yes | 3 |
| intercom.com | yes | yes | no | no | yes | 3 |
| lattice.com | no | yes | yes | no | yes | 3 |
| n8n.io | yes | yes | no | no | yes | 3 |
| planetscale.com | yes | yes | no | no | yes | 3 |
| qdrant.tech | no | yes | yes | no | yes | 3 |
| screen.studio | no | yes | yes | no | yes | 3 |
| secureframe.com | yes | yes | no | no | yes | 3 |
| sentry.io | yes | yes | no | no | yes | 3 |
| webflow.com | no | yes | yes | no | yes | 3 |
| activepieces.com | yes | yes | no | no | no | 2 |
| appsmith.com | yes | yes | no | no | no | 2 |
| ashbyhq.com | yes | yes | no | no | no | 2 |
| baseten.co | yes | no | no | no | yes | 2 |
| browserbase.com | yes | no | no | no | yes | 2 |
| cal.com | yes | no | no | no | yes | 2 |
| checkr.com | yes | yes | no | no | no | 2 |
| clickup.com | no | yes | no | no | yes | 2 |
| convex.dev | yes | no | no | no | yes | 2 |
| cultureamp.com | no | yes | no | no | yes | 2 |
| dagster.io | no | yes | no | no | yes | 2 |
| dub.co | no | yes | yes | no | no | 2 |
| e2b.dev | yes | no | no | no | yes | 2 |
| fillout.com | yes | no | no | no | yes | 2 |
| firehydrant.io | yes | no | no | no | yes | 2 |
| fly.io | yes | no | no | no | yes | 2 |
| gem.com | no | yes | no | no | yes | 2 |
| greenhouse.io | no | yes | yes | no | no | 2 |
| hightouch.com | yes | no | no | no | yes | 2 |
| launchdarkly.com | yes | yes | no | no | no | 2 |
| linear.app | yes | no | no | no | yes | 2 |
| liveblocks.io | yes | no | no | no | yes | 2 |
| mercury.com | yes | yes | no | no | no | 2 |
| modal.com | no | no | no | yes | yes | 2 |
| neon.tech | yes | no | no | no | yes | 2 |
| observablehq.com | yes | no | no | no | yes | 2 |
| posthog.com | yes | no | no | no | yes | 2 |
| remote.com | yes | yes | no | no | no | 2 |
| replicate.com | yes | no | no | no | yes | 2 |
| snowplow.io | no | yes | no | no | yes | 2 |
| stytch.com | yes | no | no | no | yes | 2 |
| supabase.com | no | no | no | yes | yes | 2 |
| turso.tech | no | yes | no | no | yes | 2 |
| upstash.com | yes | no | no | no | yes | 2 |
| wandb.ai | yes | yes | no | no | no | 2 |
| workos.com | no | yes | no | no | yes | 2 |
| 1password.com | yes | no | no | no | no | 1 |
| airtable.com | yes | no | no | no | no | 1 |
| alloy.com | no | yes | no | no | no | 1 |
| canny.io | no | no | no | no | yes | 1 |
| coda.io | yes | no | no | no | no | 1 |
| deepnote.com | no | yes | no | no | no | 1 |
| documenso.com | yes | no | no | no | no | 1 |
| getdbt.com | yes | no | no | no | no | 1 |
| heap.io | yes | no | no | no | no | 1 |
| logrocket.com | yes | no | no | no | no | 1 |
| miro.com | yes | no | no | no | no | 1 |
| slite.com | yes | no | no | no | no | 1 |
| tooljet.com | no | no | no | no | yes | 1 |
| clockwise.com | no | no | no | no | no | 0 |
| drata.com | no | no | no | no | no | 0 |
| gusto.com | no | no | no | no | no | 0 |
| highlight.io | no | no | no | no | no | 0 |
| missiveapp.com | no | no | no | no | no | 0 |
| ramp.com | no | no | no | no | no | 0 |
| reclaim.ai | no | no | no | no | no | 0 |
| scribehow.com | no | no | no | no | no | 0 |
| withpersona.com | no | no | no | no | no | 0 |
Methodology
Each site was requested once over HTTPS with a declared research user-agent, following redirects. The homepage, /robots.txt and /llms.txt were fetched and parsed with deterministic rules. No paid SEO provider, LLM, analytics or Search Console data was used. Every host was probed beforehand to confirm it answers; hosts that refuse a research user-agent remain in the sample and are recorded as failed fetches rather than dropped.
Requests used the user-agent RadiusRankResearchBot/1.0 (+https://www.radiusrank.com/research; research@radiusrank.com), followed redirects, and timed out after 12 seconds. Each site received exactly three requests. The generator is committed at scripts/generate-saas-readiness-study.mjs, so the run is reproducible against the same list.
Limitations
- The sample is a named, non-random selection of YC, Product Hunt and venture-funded B2B SaaS. It is not a random draw from all SaaS, and it deliberately excludes enterprise brands, whose readings say more about SEO headcount than about the category.
- Homepage only. A site may expose schema on pricing, docs or product pages that this study did not fetch.
- 4 of 100 sites refused the research user-agent. They stay in the denominator, so their signals count as absent rather than being dropped.
- Naming an AI crawler is recorded as a deliberate choice, not as a good one. A site that disallows every agent is counted the same as a site that allows them.
- The study measures whether a signal is present, not whether an answer engine acted on it. It cannot show causation with citations or traffic.
- RadiusRank sells AI SEO services and has a commercial interest in the category this study measures.
Data
The full record set for all 100 sites, including per-site HTTP status, the schema types found and homepage hashes, is published as saas-ai-readiness-2026-08-14.json.