RadiusRankResearch

Original RadiusRank research

91 of 100 SaaS sites point robots.txt at a sitemap. 15 name an AI crawler.

We fetched the homepage, robots.txt and llms.txt of 100 YC-backed, Product Hunt launched and venture-funded B2B SaaS companies on August 14, 2026. They are close to universal on the conventions search engines have rewarded for twenty years — 98 carry a title, 91 point robots.txt at a sitemap — and close to absent on the ones answer engines read. Only 3 of 100 carried all five signals this study measures.

Updated 100-site dataset, every company named

Key findings

  • The gap is not effort, it is vintage. 91 of 100 sites point robots.txt at a sitemap; only 15 name a single AI crawler. These teams maintain their robots file. They have simply never added the agents that read it now.
  • 23 of 100 (23%) declare themselves as software. Every company in this sample sells a product or an app, but only this many say so in a machine-readable way, even though 60 manage Organization schema.
  • Only 3 of 100 carried all five signalsinngest.com, render.com, weaviate.io.
  • llms.txt is no longer the rare one. 67 of 100 served a valid file, which is more than managed Organization schema. The convention arrived faster in this cohort than the older ones did.

Signal by signal

SignalSitesShare
Homepage returned to a research crawler96 / 10096%
Title tag98 / 10098%
Meta description77 / 10077%
robots.txt points at a sitemap91 / 10091%
Any JSON-LD at all66 / 10066%
Canonical URL68 / 10068%
Organization schema60 / 10060%
Product or app schema23 / 10023%
Named an AI crawler15 / 10015%
Valid llms.txt67 / 10067%
FAQ schema6 / 1006%

How many signals per site

An average would hide the shape. Scored against the five signals above, the sample clusters low: most sites carry two.

  • 0 of 59
  • 1 of 513
  • 2 of 536
  • 3 of 523
  • 4 of 516
  • 5 of 53

Which AI agents get named

Across the 15 sites that named any agent, these are the ones that appear. A site naming an agent has made a deliberate choice about it, whether that choice is allow or disallow.

AgentSites naming it
PerplexityBot12
GPTBot11
ClaudeBot11
ChatGPT-User10
CCBot10
OAI-SearchBot8
Google-Extended8
Bytespider6
Applebot-Extended5
Perplexity-User3
meta-externalagent3
Claude-SearchBot3
Claude-User2

All 100 sites

Ordered by how many of the five signals each carries. Every company in the sample is named, because a benchmark you cannot audit is not evidence.

CompanyCanonical URLOrganization schemaProduct or app schemaNamed an AI crawlerValid llms.txtScore
inngest.comyesyesyesyesyes5
render.comyesyesyesyesyes5
weaviate.ioyesyesyesyesyes5
airbyte.comyesyesyesnoyes4
amplitude.comyesyesnoyesyes4
atlan.comyesyesnoyesyes4
attio.comnoyesyesyesyes4
calendly.comyesyesnoyesyes4
clerk.comyesyesnoyesyes4
folk.appnoyesyesyesyes4
guidde.comyesyesyesyesno4
knock.appyesyesyesnoyes4
novu.coyesyesyesnoyes4
railway.appyesyesyesnoyes4
resend.comyesyesyesnoyes4
rootly.comnoyesyesyesyes4
tally.soyesyesyesnoyes4
temporal.ioyesyesyesnoyes4
vercel.comyesyesyesnoyes4
arcade.softwarenoyesyesnoyes3
astronomer.ioyesyesnonoyes3
brex.comyesyesnonoyes3
budibase.comyesyesnonoyes3
chronosphere.ioyesyesnonoyes3
comet.comyesyesnonoyes3
deel.comyesnonoyesyes3
figma.comyesyesnoyesno3
framer.comyesyesnonoyes3
front.comyesyesnonoyes3
fullstory.comyesyesnonoyes3
grafana.comyesyesnonoyes3
hex.techyesyesyesnono3
incident.ioyesyesnonoyes3
intercom.comyesyesnonoyes3
lattice.comnoyesyesnoyes3
n8n.ioyesyesnonoyes3
planetscale.comyesyesnonoyes3
qdrant.technoyesyesnoyes3
screen.studionoyesyesnoyes3
secureframe.comyesyesnonoyes3
sentry.ioyesyesnonoyes3
webflow.comnoyesyesnoyes3
activepieces.comyesyesnonono2
appsmith.comyesyesnonono2
ashbyhq.comyesyesnonono2
baseten.coyesnononoyes2
browserbase.comyesnononoyes2
cal.comyesnononoyes2
checkr.comyesyesnonono2
clickup.comnoyesnonoyes2
convex.devyesnononoyes2
cultureamp.comnoyesnonoyes2
dagster.ionoyesnonoyes2
dub.conoyesyesnono2
e2b.devyesnononoyes2
fillout.comyesnononoyes2
firehydrant.ioyesnononoyes2
fly.ioyesnononoyes2
gem.comnoyesnonoyes2
greenhouse.ionoyesyesnono2
hightouch.comyesnononoyes2
launchdarkly.comyesyesnonono2
linear.appyesnononoyes2
liveblocks.ioyesnononoyes2
mercury.comyesyesnonono2
modal.comnononoyesyes2
neon.techyesnononoyes2
observablehq.comyesnononoyes2
posthog.comyesnononoyes2
remote.comyesyesnonono2
replicate.comyesnononoyes2
snowplow.ionoyesnonoyes2
stytch.comyesnononoyes2
supabase.comnononoyesyes2
turso.technoyesnonoyes2
upstash.comyesnononoyes2
wandb.aiyesyesnonono2
workos.comnoyesnonoyes2
1password.comyesnononono1
airtable.comyesnononono1
alloy.comnoyesnonono1
canny.iononononoyes1
coda.ioyesnononono1
deepnote.comnoyesnonono1
documenso.comyesnononono1
getdbt.comyesnononono1
heap.ioyesnononono1
logrocket.comyesnononono1
miro.comyesnononono1
slite.comyesnononono1
tooljet.comnonononoyes1
clockwise.comnonononono0
drata.comnonononono0
gusto.comnonononono0
highlight.iononononono0
missiveapp.comnonononono0
ramp.comnonononono0
reclaim.ainonononono0
scribehow.comnonononono0
withpersona.comnonononono0

Methodology

Each site was requested once over HTTPS with a declared research user-agent, following redirects. The homepage, /robots.txt and /llms.txt were fetched and parsed with deterministic rules. No paid SEO provider, LLM, analytics or Search Console data was used. Every host was probed beforehand to confirm it answers; hosts that refuse a research user-agent remain in the sample and are recorded as failed fetches rather than dropped.

Requests used the user-agent RadiusRankResearchBot/1.0 (+https://www.radiusrank.com/research; research@radiusrank.com), followed redirects, and timed out after 12 seconds. Each site received exactly three requests. The generator is committed at scripts/generate-saas-readiness-study.mjs, so the run is reproducible against the same list.

Limitations

  • The sample is a named, non-random selection of YC, Product Hunt and venture-funded B2B SaaS. It is not a random draw from all SaaS, and it deliberately excludes enterprise brands, whose readings say more about SEO headcount than about the category.
  • Homepage only. A site may expose schema on pricing, docs or product pages that this study did not fetch.
  • 4 of 100 sites refused the research user-agent. They stay in the denominator, so their signals count as absent rather than being dropped.
  • Naming an AI crawler is recorded as a deliberate choice, not as a good one. A site that disallows every agent is counted the same as a site that allows them.
  • The study measures whether a signal is present, not whether an answer engine acted on it. It cannot show causation with citations or traffic.
  • RadiusRank sells AI SEO services and has a commercial interest in the category this study measures.

Data

The full record set for all 100 sites, including per-site HTTP status, the schema types found and homepage hashes, is published as saas-ai-readiness-2026-08-14.json.

Start with the evidence

Run these checks on your own SaaS site.

Every signal here is deterministic and public. You can verify each one on your own domain before deciding what to fix.

Run the free diagnostic