Fish Audio raises $52 million to challenge ElevenLabs in the AI voice arms race Fish Audio, a San Francisco startup founded by former Nvidia researcher Shijia Liao, has raised a $52 million seed round led by Coreline Ventures and Capital Today, reaching $21 million in annual recurring revenue and 8 million users before closing its first institutional check. The company launched its S2.1 Pro model supporting real-time text-to-speech, voice cloning, and voice agents across 83 languages, with 67% of listeners preferring its outputs over competitors in blind tests. The funding positions Fish Audio to challenge ElevenLabs in the AI voice market, particularly in the creator economy and developer use cases. The San Francisco startup hit $21 million in annual recurring revenue and 8 million users before closing its first institutional check, suggesting the AI voice market still has room for a serious second act. Fish Audio didn't start with a pitch deck. It started with Shijia Liao, a former Nvidia video researcher and self-described VTuber fan, training a voice model on a single gaming GPU in his bedroom because every synthetic voice he could find sounded hollow. That open-source project, Fish Speech, became one of the most-starred voice repositories on GitHub. A year later, as TechCrunch reported on July 28, the company has raised a $52 million seed round led by Coreline Ventures and Capital Today, with participation from 645 Ventures, HF0, 359 Capital, and several others. The number is striking for a seed. It is not unheard of in AI infrastructure, where the cost of training and the speed of the market reward early scale, but it puts Fish Audio in rarefied company for a business still in its first year of commercial life. The round comes after the company reached $21 million in annual recurring revenue and more than 8 million users across creators, developers, and enterprise customers - none of it backed by institutional capital until now. The funding announcement arrives alongside the launch of Fish Audio's S2.1 Pro model, which supports real-time text-to-speech, voice cloning, and voice agents across 83 languages. The company says 67% of listeners preferred its voice outputs over those of competing models in blind listening tests. According to Artificial Analysis, S2.1 Pro earned an Elo score of 1,153, placing it 13th on the AI audio leaderboard, with particular strengths in low-latency streaming and natural language control over emotion and prosody. Fish Audio plans to make S2.1 Pro available free to every developer via API starting in late August. It's free, at least for developers. Free API access as a growth lever is a smart bet in infrastructure, and Fish Audio has precedent on its side. The open-source Fish Speech release is what built the company's initial user base without a dollar of marketing spend. The question is whether that same flywheel works at enterprise scale, where ElevenLabs has already locked in clients including Deutsche Telekom, Boston Consulting Group, and Revolut, and claims usage by 41% of Fortune 500 companies. The gap Fish Audio is actually betting against ElevenLabs is the obvious comparison, but the numbers show how steep the climb is. In February 2026, ElevenLabs raised $500 million from Sequoia Capital and Andreessen Horowitz at an $11 billion valuation, as CNBC reported. By May the company had hit $500 million in ARR and was reportedly in early talks for a secondary share sale at a $22 billion valuation. That is a category winner, not just a first mover. Cartesia, the other credible rival, raised a $100 million Series B in October 2025, and was founded by Albert Gu, one of the researchers behind the Mamba state-space architecture that underpins many of the fastest low-latency voice models. These are not small incumbents. So where does Fish Audio fit? The honest answer is: probably not in the same boardroom as ElevenLabs, at least not yet. The more interesting opportunity is in the creator economy and the long tail of developer use cases that a $22 billion company has less incentive to serve cheaply. Fish Audio grew to 8 million users on the back of an anime-adjacent, VTuber-obsessed community that cares deeply about expressive voice cloning in Japanese, Korean, and Mandarin. That is not a fringe market. It is a real one that the voice AI incumbents have historically treated as secondary. The 83-language coverage and the self-hosting option for teams that want full data control are also pointed choices. Enterprise buyers in regulated industries - healthcare, finance, legal - often cannot send audio through a shared cloud API. A model you can run on your own infrastructure is a different product from a subscription to someone else's platform, and it is a gap that neither ElevenLabs nor Cartesia has fully closed. Fish Audio's backers are betting the market is big enough that a second serious player can reach real scale, and the ARR trajectory suggests they are not wrong to try. Whether $52 million is enough to sustain the model training and infrastructure costs while the category consolidates is the harder question. The reference point is not encouraging: in AI voice, as in most infrastructure markets, the gap between first and third closes faster than founders expect. Liao built something real. Now the question is whether a bedroom project that went viral can turn into a company that wins deals. Also read: Anthropic's Claude found real flaws in encryption used by billions of devices https://startupfortune.com/anthropics-claude-found-real-flaws-in-encryption-used-by-billions-of-devices/ • America's biggest power grid just told data centers it can cut their power off https://startupfortune.com/americas-biggest-power-grid-just-told-data-centers-it-can-cut-their-power-off/ • Richard Socher's Recursive Superintelligence signs $410 million AWS compute deal and calls it its smallest ever https://startupfortune.com/richard-sochers-recursive-superintelligence-signs-410-million-aws-compute-deal-and-calls-it-its-smallest-ever/