cd /news/ai-tools/a-quarter-of-yc-fall-2025-startups-a… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-95962] src=readablebyai.com β†— pub= topic=ai-tools verified=true sentiment=Β· neutral

A quarter of YC Fall 2025 startups are blank pages to AI crawlers

A ReadableByAI Index found that 25.5% of 145 YC Fall 2025 startup homepages are empty client-rendered shells to AI crawlers like GPTBot, ClaudeBot, and PerplexityBot, compared to 2.9% of 455 established SaaS companies, an 8.9x gap, as measured on August 8, 2026. The index also reported that 76.6% of YC startups have a robots.txt file and 33.1% have an llms.txt, versus 98.0% and 54.3% for established SaaS firms.

read14 min views1 publishedAug 13, 2026

THE READABLEBYAI INDEX

25.5% of 145 YC Fall 2025 homepages ship an empty client-rendered shell to GPTBot, ClaudeBot and PerplexityBot β€” no headline, no product description, nothing but a <div id="root">. Among 455 established SaaS companies measured the same way, it's 2.9%. That's a 8.9x gap. Measured 8 August 2026.

What was measured, and what wasn't #

Every homepage in this Index was fetched twice: once with a baseline browser user-agent, and once each with the twelve crawler user-agents (eleven AI-vendor crawlers plus bingbot as a non-AI control) ReadableByAI probes, all from the same datacenter IP. The readability numbers below β€” visible word count, rendering classification, robots.txt contents β€” come from the baseline fetch, and they are identity-independent: whether a page ships its content in raw HTML or hides it behind a client-side render is true for every visitor, bot or browser, and needs no further confirmation.

What we deliberately did not publish is bot-specific access results β€” which user-agents got a 200, a 403, or a challenge page. A datacenter IP claiming to be ClaudeBot is an unverified probe, not the real crawler; vendors authenticate their bots by published IP range, so a challenge to our probe cannot distinguish a genuine block from a WAF simply doubting an impostor. Only the site's own server logs settle that, and we don't have them. So this Index reports what any fetch can prove β€” content and permission β€” and leaves access claims out entirely.

One further limit, stated plainly because it is the strongest objection to this method: a site can serve different HTML to different requesters. The common implementation keys on the user-agent string, and this Index detects it β€” every company named below returned the same content to every crawler identity tested, verified by comparing the response returned to each identity against the baseline (four were byte-for-byte identical across all thirteen fetches; the other two varied by under one percent in size, consistent with per-request timestamps). The four companies that did serve crawlers a materially different page are reported separately rather than counted as failures. What an outside probe cannot rule out is a site that varies its content by verified crawler IP range rather than by user-agent. That is rare, and nothing in this dataset suggests it, but it cannot be disproven from outside β€” which is the honest reason server logs matter and probes alone are never the last word.

That limit is not a flaw in this Index; it is the boundary of what any outside probe can establish, ours included. Resolving it needs the one record we do not have: the target's own server logs, with every hit checked against each vendor's published crawler IP ranges. That separates verified crawlers from impostors, shows which pages they actually retrieved, and surfaces the fetches they abandoned midway β€” none of which is visible from the outside, and all of which is what the paid audit reads. A roadmap item, not yet built and therefore not sold: dual-origin probing from residential and datacenter networks simultaneously, which would expose IP-sensitive bot management without logs.

The comparison #

145 of 147 YC Fall 2025 companies and 455 of 514 established SaaS companies returned a clean 200 to the baseline fetch and are counted below. The rest are excluded from every percentage on this page, not silently dropped: 2 YC domains (2 non-200 responses) and 59 SaaS domains (51 non-200, 8 unreachable).

Population Clean baseline CSR_SHELL SSR_THIN SSR_FULL Median words robots.txt llms.txt
YC Fall 2025 145 / 147 37 (25.5%) 14 (9.7%) 94 (64.8%) 630 111 (76.6%) 48 (33.1%)
Established SaaS 455 / 514 13 (2.9%) 11 (2.4%) 431 (94.7%) 1239 446 (98.0%) 247 (54.3%)

CSR_SHELL = under 150 visible words in raw HTML. SSR_THIN = 150–399. SSR_FULL = 400+. Median visible words: 630 for YC vs. 1239 for SaaS β€” the typical SaaS homepage ships roughly double the readable content of the typical YC Fall 2025 homepage, before either one is judged by whether it's reachable at all.

The comparison isn't one-directional. 4 companies in the established-SaaS set serve AI crawlers more content than they served our baseline fetch β€” bot-aware, dynamic rendering that detects a non-browser request and responds with fuller markup instead of a thinner one. The two largest gaps we measured: a payroll platform serving roughly 220x the byte size to a crawler that it serves to a plain browser fetch, and an ML-ops platform at roughly 141x (both public companies; unnamed here because per-company detail belongs to the named-examples policy above). Some teams have already solved this deliberately β€” it's evidence this is a solvable engineering problem, not an unavoidable one.

The 145 YC Fall 2025 homepages, anonymized #

These are the 145 Y Combinator Fall 2025 companies that returned a clean baseline fetch, published anonymized. The distribution is the finding, not any single row in it: 25.5% of a whole recent YC batch under 150 visible words is the headline of this Index. Individual companies aren't named here, for the same reason the SaaS roster isn't published in full β€” the point of this Index is to get pages fixed, not to build a wall of names. If your company is in this batch, the free scan on the ReadableByAI homepage checks your own homepage the same way we checked this one, so you can find out where you stand without waiting on us to reach out.

The yc-001…yc-145 ids are stable within this dataset β€” the same company keeps the same id release over release, so a re-verified row can be tracked across waves β€” but they carry no mapping we publish. There is no lookup table connecting an id back to a domain anywhere on this site.

Anonymized ID Visible words Classification
yc-001 1 CSR_SHELL
yc-002 1 CSR_SHELL
yc-003 1 CSR_SHELL
yc-004 2 CSR_SHELL
yc-005 5 CSR_SHELL
yc-006 5 CSR_SHELL
yc-007 5 CSR_SHELL
yc-008 6 CSR_SHELL
yc-009 6 CSR_SHELL
yc-010 6 CSR_SHELL
yc-011 6 CSR_SHELL
yc-012 7 CSR_SHELL
yc-013 7 CSR_SHELL
yc-014 8 CSR_SHELL
yc-015 8 CSR_SHELL
yc-016 9 CSR_SHELL
yc-017 9 CSR_SHELL
yc-018 9 CSR_SHELL
yc-019 9 CSR_SHELL
yc-020 10 CSR_SHELL
yc-021 12 CSR_SHELL
yc-022 14 CSR_SHELL
yc-023 15 CSR_SHELL
yc-024 18 CSR_SHELL
yc-025 19 CSR_SHELL
yc-026 20 CSR_SHELL
yc-027 27 CSR_SHELL
yc-028 29 CSR_SHELL
yc-029 30 CSR_SHELL
yc-030 32 CSR_SHELL
yc-031 36 CSR_SHELL
yc-032 41 CSR_SHELL
yc-033 52 CSR_SHELL
yc-034 67 CSR_SHELL
yc-035 88 CSR_SHELL
yc-036 93 CSR_SHELL
yc-037 130 CSR_SHELL
yc-038 166 SSR_THIN
yc-039 167 SSR_THIN
yc-040 189 SSR_THIN
yc-041 212 SSR_THIN
yc-042 223 SSR_THIN
yc-043 304 SSR_THIN
yc-044 309 SSR_THIN
yc-045 314 SSR_THIN
yc-046 314 SSR_THIN
yc-047 369 SSR_THIN
yc-048 386 SSR_THIN
yc-049 388 SSR_THIN
yc-050 391 SSR_THIN
yc-051 395 SSR_THIN
yc-052 401 SSR_FULL
yc-053 440 SSR_FULL
yc-054 442 SSR_FULL
yc-055 454 SSR_FULL
yc-056 454 SSR_FULL
yc-057 515 SSR_FULL
yc-058 519 SSR_FULL
yc-059 522 SSR_FULL
yc-060 523 SSR_FULL
yc-061 534 SSR_FULL
yc-062 539 SSR_FULL
yc-063 540 SSR_FULL
yc-064 541 SSR_FULL
yc-065 554 SSR_FULL
yc-066 568 SSR_FULL
yc-067 574 SSR_FULL
yc-068 585 SSR_FULL
yc-069 596 SSR_FULL
yc-070 607 SSR_FULL
yc-071 612 SSR_FULL
yc-072 629 SSR_FULL
yc-073 630 SSR_FULL
yc-074 638 SSR_FULL
yc-075 644 SSR_FULL
yc-076 651 SSR_FULL
yc-077 651 SSR_FULL
yc-078 652 SSR_FULL
yc-079 655 SSR_FULL
yc-080 659 SSR_FULL
yc-081 690 SSR_FULL
yc-082 691 SSR_FULL
yc-083 706 SSR_FULL
yc-084 728 SSR_FULL
yc-085 729 SSR_FULL
yc-086 742 SSR_FULL
yc-087 748 SSR_FULL
yc-088 757 SSR_FULL
yc-089 775 SSR_FULL
yc-090 777 SSR_FULL
yc-091 786 SSR_FULL
yc-092 796 SSR_FULL
yc-093 835 SSR_FULL
yc-094 850 SSR_FULL
yc-095 876 SSR_FULL
yc-096 906 SSR_FULL
yc-097 908 SSR_FULL
yc-098 941 SSR_FULL
yc-099 943 SSR_FULL
yc-100 954 SSR_FULL
yc-101 962 SSR_FULL
yc-102 976 SSR_FULL
yc-103 981 SSR_FULL
yc-104 1003 SSR_FULL
yc-105 1049 SSR_FULL
yc-106 1065 SSR_FULL
yc-107 1069 SSR_FULL
yc-108 1096 SSR_FULL
yc-109 1101 SSR_FULL
yc-110 1119 SSR_FULL
yc-111 1154 SSR_FULL
yc-112 1160 SSR_FULL
yc-113 1169 SSR_FULL
yc-114 1187 SSR_FULL
yc-115 1188 SSR_FULL
yc-116 1213 SSR_FULL
yc-117 1216 SSR_FULL
yc-118 1279 SSR_FULL
yc-119 1292 SSR_FULL
yc-120 1323 SSR_FULL
yc-121 1368 SSR_FULL
yc-122 1377 SSR_FULL
yc-123 1420 SSR_FULL
yc-124 1423 SSR_FULL
yc-125 1509 SSR_FULL
yc-126 1522 SSR_FULL
yc-127 1533 SSR_FULL
yc-128 1538 SSR_FULL
yc-129 1566 SSR_FULL
yc-130 1589 SSR_FULL
yc-131 1615 SSR_FULL
yc-132 1635 SSR_FULL
yc-133 1641 SSR_FULL
yc-134 1642 SSR_FULL
yc-135 1656 SSR_FULL
yc-136 1689 SSR_FULL
yc-137 1748 SSR_FULL
yc-138 1756 SSR_FULL
yc-139 1812 SSR_FULL
yc-140 1943 SSR_FULL
yc-141 2198 SSR_FULL
yc-142 2299 SSR_FULL
yc-143 2895 SSR_FULL
yc-144 3868 SSR_FULL
yc-145 3952 SSR_FULL

Six companies, named #

This Index doesn't publish the roster of 600 domains it measured β€” see below for why. What it can publish, without that tradeoff, are results any reader can reproduce in ten seconds: six large, well-resourced, publicly recognizable companies, each checked exactly the way every other homepage in this dataset was checked.

These aren't obscure or under-resourced teams β€” that's the point. A company with Palantir's or Duolingo's engineering headcount isn't serving zero words to AI crawlers because it can't afford server-side rendering. It's serving zero words because a popular frontend framework's default configuration ships a client-rendered shell unless someone deliberately turns on server rendering, and nobody happened to. This is a framework-default problem, not a competence problem β€” and none of the six disallows any AI crawler in robots.txt, which is what you'd expect to see if this were policy instead of an accident. A page that's deliberately kept off-limits to GPTBot says so in robots.txt; a page that's just empty says nothing, because nobody meant for it to be empty.

Company Visible words Classification
duolingo.com 1 CSR_SHELL
palantir.com 4 CSR_SHELL
qualys.com 8 CSR_SHELL
grindr.com 118 CSR_SHELL
blackberry.com 216 SSR_THIN
substack.com 284 SSR_THIN

Reproduce any of these yourself β€” raw HTML, JavaScript disabled, visible word count only:

curl -s https://duolingo.com | python3 -c "import re,sys; h=sys.stdin.read(); h=re.sub(r'<(script|style)\[^>]*>.*?</\1>','',h,flags=re.S|re.I); h=re.sub(r'<[^>]+>',' ',h); print(len(h.split()))"

What we're not publishing, and why #

Earlier versions of this page listed all 600 measured domains. We took that table down. This dataset doubles as a private outreach list β€” companies we contact directly when we find their homepage invisible to AI crawlers β€” and publishing the full roster would trade a company's incentive to fix the problem for our incentive to publish a bigger list. We'd rather have the fix. Beyond the six SaaS companies named above: sixteen other companies in this dataset had the same finding. We are notifying each of them privately rather than publishing a wall of names. That's not concealment β€” it's the actual policy, stated plainly: the point of this Index is to get pages fixed, not a leaderboard of who's been caught. If you run a company and want to know where you stand, the free scan on the ReadableByAI homepage checks your own site the same way, right now, without waiting for us to reach out.

Method and reproducibility #

Every homepage was fetched with a plain HTTP GET β€” no headless browser, no JavaScript execution β€” because that's what GPTBot, ClaudeBot and PerplexityBot do. Visible word count is text content extracted from the raw HTML response, minus script and style contents. The three classification bands are fixed thresholds: CSR_SHELL under 150 visible words, SSR_THIN 150–399, SSR_FULL 400 or more. robots.txt is read directly and reported as a literal fact β€” present or absent, nothing inferred about intent. The engine behind this Index is open source at github.com/abouchard11/geo-crawl-audit. Anyone can clone it, point it at their own list, and check our numbers or produce their own.

Fixed it? We'll re-verify free. #

If your homepage came back CSR_SHELL or SSR_THIN when we measured it β€” named above or notified privately β€” and you've since shipped server-rendered content, tell us and we'll re-probe it at no charge β€” no catch, no upsell attached. The goal of this Index is an accurate public record, not a leaderboard anyone stays stuck on. Tell us: alex+reverify@midnightdev.dev.

Named here? Talk to us directly. #

If your company appears in this report β€” named in the table above or described anonymously β€” this is your direct line: alex+named@midnightdev.dev. A real person answers, same week. If we didn't have a working public contact channel for your company before publication, this address is that channel β€” no form, no gatekeeper. Disputes and reproduction mismatches follow the process on the Corrections page; re-verification after a fix is free, always.

Four case files go deeper on the findings behind this Index, and the free scan on the ReadableByAI homepage checks your own site the same way.

Invisible without JavaScript

What a client-rendered shell looks like to a crawler that never runs the script β€” the raw HTML, side by side with what a browser paints.

Startups are nearly 9x more invisible

The YC Fall 2025 vs. established-SaaS comparison behind this Index, in full β€” including the two findings that didn't survive review.

A menu on a locked door

Why publishing an llms.txt file does nothing if the homepage behind it is a blank shell β€” permission without content to permit.

β€œWhy they block GPTBot” is usually wrong

The difference between a robots.txt disallow, a WAF challenge, and an empty page β€” and why only one of the three is in this dataset.

This Index is measured periodically, not once. The next wave will show what changed β€” which domains fixed a CSR_SHELL homepage, which didn't, and whether the gap between YC Fall 2025 and established SaaS narrowed or widened.

── more in #ai-tools 4 stories Β· sorted by recency
── more on @readablebyai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/a-quarter-of-yc-fall…] indexed:0 read:14min 2026-08-13 Β· β€”