- Fish Audio raised $52M in seed funding led by Coreline Ventures and Capital Today, with participation from 645 Ventures, HF0, and others [1] - The company hit $21M ARR and 8M+ users within its first year, with enterprise customers including OpenAI, HeyGen, and LiveKit [2] - Its flagship S2.1 Pro model was preferred by 67% of listeners over competitors in blind tests and supports 83 languages [2] - Funds will go toward voice-native LLMs, speech-to-speech translation, and an enterprise sales team
[[3]](https://siliconangle.com/2026/07/28/fish-audio-makes-splash-raising-52m-seed-funding-ai-voices/) - Fish Audio competes with ElevenLabs, which raised $500M at an $11B valuation earlier this year
[[3]](https://siliconangle.com/2026/07/28/fish-audio-makes-splash-raising-52m-seed-funding-ai-voices/)
Fish Audio, operating under the legal name Hanabi AI Inc., has raised $52 million in seed funding as the AI voice startup marks its first anniversary with $21 million in annual recurring revenue and more than 8 million users. Coreline Ventures and Capital Today co-led the round, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners [1].
The Palo Alto-based company builds AI voice models for real-time text-to-speech, voice cloning, and voice agents. Its platform can clone a voice from a five-second audio clip in roughly 15 seconds and supports 83 languages with word-level emotion controls driven by more than 15,000 natural language prompts [2].
Fish Audio's customer roster already includes several prominent AI companies: OpenAI, HeyGen, Retell, LiveKit, Telnyx, and Sanas all use its technology. The startup plans to use the new capital to move beyond text-to-speech into voice-native large language models and real-time speech-to-speech translation, while also building out its enterprise sales team [1].
The Round #
The $52 million seed is unusually large for a pre-Series A company and reflects the rapid commercialization Fish Audio achieved in its first year. The round was co-led by Coreline Ventures and Capital Today, with a syndicate that includes 645 Ventures, HF0, Parable, Carya Venture Partners, and Alphalist Partners, plus unnamed angel investors [1].
No valuation was disclosed. The company said it will deploy the capital across three priorities: expanding its model lineup with voice-native LLMs and speech-to-speech tools, deepening developer integrations with partners like LiveKit and Retell AI, and hiring an enterprise sales team [2].
The Founders and Origin Story #
Fish Audio was co-founded by Chief Scientist Shijia Liao and CEO Rissa Cao. Liao, a former Nvidia video researcher and longtime VTuber and anime enthusiast, started training voice models on a single gaming GPU out of frustration with the monotonous quality of synthetic voices at the time [1].
His open-source Fish Speech project on GitHub accumulated more than 31,000 stars, attracting indie developers, video game designers, and content creators before the pair formalized the effort into a venture-backed company [2].
The Product #
Fish Audio's flagship model, S2.1 Pro, was preferred by 67% of listeners over competitors in blind listening tests, according to the company. The platform supports voice cloning from five-second audio samples, real-time text-to-speech across 83 languages, and fine-grained emotion control at the word level [2].
For enterprise customers, the company offers on-premises deployment and HIPAA-compliant configurations. Fish Audio said it will make S2.1 Pro available for free starting at the end of August [3].
Competitive Landscape #
Fish Audio enters a crowded AI voice market headlined by ElevenLabs, which raised $500 million at an $11 billion valuation earlier this year. Other competitors include Amazon's Polly, Google's Cloud Text-to-Speech, and a growing cohort of startups targeting real-time voice synthesis [3].
Fish Audio's open-source roots and developer community give it a different distribution channel than most competitors. The 31,000-star GitHub repository serves as a funnel for both individual creators and enterprise prospects who first experiment with the open-source models before adopting the hosted commercial platform [2].
What's Next #
The company's roadmap extends well beyond text-to-speech. Fish Audio plans to build voice-native large language models — AI systems designed from the ground up to process and generate speech, rather than treating voice as an add-on to text-based models [1].
It is also developing real-time speech-to-speech translation tools and expanding API integrations through partnerships with voice infrastructure companies like LiveKit and Retell AI. The enterprise sales team buildout signals a shift toward larger contract sizes as the company moves upmarket from its creator and developer base [2].
Further sources #
The stories that matter, in one email. Free — unsubscribe anytime.