# I ran cosine similarity on every YC startup. 2025 batches are 2.5x more repetitive than 2015

> Source: <https://dev.to/jakobgreenfeld/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-25x-more-repetitive-than-2015-354n>
> Published: 2026-09-30 07:32:13+00:00

Everyone says every YC startup is now the same AI agent. I ran cosine similarity on all 6,142 of them to see whether that is true, and by how much. In 2015, one in five companies in a batch had a near twin among earlier YC companies. In 2025 it was two in five.

6,142 Y Combinator companies with a live domain, 48 batches from Summer 2005 through Fall 2026. Each one got a vector from text-embedding-3-large at 1024 dimensions.

The catch: I did not embed homepages. Homepage copy varies more by who wrote it than by what the company does, and an embedding of raw copy mostly measures writing style. Instead, gpt-5-mini wrote a five-sentence summary of each company from two inputs, a rendered homepage (ScrapingBee) and search snippets for the domain (Serper). The prompt forces the same five questions in the same order every time: product, buyer, delivery, market, pricing. Same shape of text for a 2008 company and a 2026 one. 206 companies had no usable homepage and fell back to YC's own description; those skew old, so they push against the result rather than for it.

Similarity is cosine between the normalised vectors.

For each company, take its best cosine score against a random 500 companies from earlier batches. Count how many in the batch clear 0.80. Repeat with ten different random 500s and average.

Full analysis and data here: [https://fundingwatcher.com/research/yc-batch-similarity/](https://fundingwatcher.com/research/yc-batch-similarity/)
