{"slug": "i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015", "title": "I ran cosine similarity on every YC startup. 2025 batches are 2.5x more repetitive than 2015", "summary": "A developer embedded all 6,142 Y Combinator companies with live domains across 48 batches (Summer 2005 through Fall 2026) using text-embedding-3-large at 1024 dimensions and found that 2025 batches are roughly 2.5x more repetitive than 2015. In 2015, one in five companies in a batch had a near twin among earlier YC companies, compared with two in five in 2025, measured by counting batch members whose best cosine similarity against a random sample of 500 earlier companies cleared 0.80. To avoid measuring writing style, the analysis embedded gpt-5-mini summaries generated from rendered homepages and search snippets rather than raw homepage copy.", "body_md": "Everyone says every YC startup is now the same AI agent. I ran cosine similarity on all 6,142 of them to see whether that is true, and by how much. In 2015, one in five companies in a batch had a near twin among earlier YC companies. In 2025 it was two in five.\n\n6,142 Y Combinator companies with a live domain, 48 batches from Summer 2005 through Fall 2026. Each one got a vector from text-embedding-3-large at 1024 dimensions.\n\nThe catch: I did not embed homepages. Homepage copy varies more by who wrote it than by what the company does, and an embedding of raw copy mostly measures writing style. Instead, gpt-5-mini wrote a five-sentence summary of each company from two inputs, a rendered homepage (ScrapingBee) and search snippets for the domain (Serper). The prompt forces the same five questions in the same order every time: product, buyer, delivery, market, pricing. Same shape of text for a 2008 company and a 2026 one. 206 companies had no usable homepage and fell back to YC's own description; those skew old, so they push against the result rather than for it.\n\nSimilarity is cosine between the normalised vectors.\n\nFor each company, take its best cosine score against a random 500 companies from earlier batches. Count how many in the batch clear 0.80. Repeat with ten different random 500s and average.\n\nFull analysis and data here: [https://fundingwatcher.com/research/yc-batch-similarity/](https://fundingwatcher.com/research/yc-batch-similarity/)", "url": "https://wpnews.pro/news/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015", "canonical_source": "https://dev.to/jakobgreenfeld/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-25x-more-repetitive-than-2015-354n", "published_at": "2026-09-30 07:32:13+00:00", "updated_at": "2026-09-30 07:46:48.763651+00:00", "lang": "en", "topics": ["ai-startups", "machine-learning", "natural-language-processing"], "entities": ["Y Combinator", "text-embedding-3-large", "gpt-5-mini", "ScrapingBee", "Serper"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015", "markdown": "https://wpnews.pro/news/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015.md", "text": "https://wpnews.pro/news/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015.txt", "jsonld": "https://wpnews.pro/news/i-ran-cosine-similarity-on-every-yc-startup-2025-batches-are-2-5x-more-than-2015.jsonld"}}