[ AI Models & Platforms
](https://www.unite.ai/series/artificial-intelligence/)
[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)
Fish Audio has raised $52 million in a seed round co-led by Coreline Ventures and Capital Today, money that arrives at a Palo Alto text-to-speech company that has spent the past year giving its models away and charging for everything around them. Fish Audio says more than 8 million people now use those models through either the open-weight releases or its hosted platform, and that the business is running at $21 million in annual recurring revenue.
The round was disclosed on July 28, 2026, with 359 Capital, the HF0 founder residency and five other funds joining, and was first reported by TechCrunch. CEO and co-founder Rissa Cao said the company had been running efficiently enough on open-source distribution and creator subscriptions that it did not need outside capital, and went out to fund more advanced models and a push into enterprise accounts. A $52 million round still carrying the seed label, at a company with eight figures of recurring revenue, marks how fast voice moved from a product feature to an infrastructure line item.
From open weights to a free API #
The company began as a side project by Shijia Liao, a former Nvidia (NVDA ) researcher who trained a speech model on a single GPU and published it. That repository, Fish Speech, now carries more than 31,000 stars and a following among indie developers and game studios. Liao is Fish Audio’s chief scientist, and five models have shipped in roughly a year: four speech generators and one speech-to-text system.
Three of the speech generators are out in the open. Fish Audio published S2’s weights on March 9, 2026, along with fine-tuning code and a streaming inference engine. S2 is a 4-billion-parameter model trained on more than 10 million hours of audio across roughly 80 languages, steered by free-form tags such as [whisper]
or [professional broadcast tone]
dropped inline at the word level. The company counts more than 15,000 of those controls.
The newest model, S2.1 Pro, is the one it holds back from open release, and even that is currently free over the API: no hard character cap, 83 languages and no service-level guarantee, with the free window extended through August 31, 2026. Paid plans carry the latency and uptime commitments, and the company asks products above $1 million in ARR to talk to it before building on the free tier. Fish Audio credits a rebuilt inference stack, including its own FP8 GPU kernel library, for making that giveaway affordable, and reports roughly 70 milliseconds to first audio on a single request.
That is the commercial logic of the business. Open weights and a free frontier model buy distribution among developers; revenue comes from the companies that need contractual latency. Fish Audio says enterprise customers including HeyGen, Livekit, Retell, Sanas and OpenArt already run on its APIs, and Cao describes demand splitting by use case: avatar products want realism, game studios want expressive character voices, and voice-agent companies want low latency that still sounds human. That last segment is the one moving fastest, as mainstream assistants add spoken interfaces — Anthropic brought its Opus and Sonnet models into Claude’s voice mode in July 2026 — and as smaller studios chase the same niche, including Tryll Engine with on-device game dialogue.
Who wrote the checks #
The two leads are an unusual pairing for a US developer-tools company. Coreline Ventures is a small cross-border fund spun out of DCM Ventures, writing early checks across the US, Japan and Korea from a deliberately compact portfolio that already includes a Japanese generative-voice lab. Capital Today is the China consumer-internet franchise Kathy Xu founded in 2005, with JD.com (JD ), Meituan, ByteDance and Moonshot AI among its holdings and an evergreen fund of roughly $2.8 billion. 359 Capital, which separated from Sapphire Ventures in November 2025 with about $300 million under management, invests in consumer and consumer-adjacent businesses. Consumer money, in other words, backing a developer API, on the theory that the community voice library is the durable asset.
Osuke Honda, Coreline’s co-founder and managing partner, put a condition on that theory. “A community-centric approach can only become a durable advantage if creators trust the platform,” he said in the same interview. “That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts.”
The library is user-supplied. Fish Audio asks people to submit their own voices for training and pays them when a voice is used, and the public library now holds more than two million voices. That model drew complaints earlier in 2026, when creators said voices had been uploaded without their consent and that takedowns moved slowly. On July 4, 2026 the company documented a dedicated voice-ownership dispute process that replaces the general support queue: claimants file from the model’s page, submit government ID or corporate authorization, and read a verification sentence aloud. Cao said removals now complete in under three minutes.
What ships next #
Two more models are on the roadmap: an audio-understanding system the company expects this year, and a speech-to-speech model still in development. Both aim at the voice-agent stack rather than one-way narration. Speech generation is already a crowded market, with ElevenLabs, Cartesia, Speechify and Krisp selling into overlapping sets of creators and developers. A raise this size is no longer the outlier it would have been two years ago; Enigma opened a $71 million seed for robot interfaces earlier in July 2026. The nearer test for Fish Audio is commercial: the free S2.1 Pro window closes at the end of August 2026, and the new capital has to convert the developers it attracted into enterprise contracts.