Fish Audio Raises $52 Million to Grow AI Voice Platform for Creators and Enterprises

Fish Audio has raised $52 million in seed funding to expand its AI voice platform, develop advanced speech models and grow its services for creators and enterprise customers.

Jul 29, 2026 - 02:59
 9
Fish Audio Raises $52 Million to Grow AI Voice Platform for Creators and Enterprises
IMAGE CREDITS: FISH AUDIO

AI voice technology startup Fish Audio has raised $52 million in seed funding as it looks to strengthen its portfolio of speech-generation models and expand its offerings for both enterprise customers and content creators.

The Palo Alto-based company announced the funding on Tuesday, saying the round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners and HF0.

Founded by former Nvidia researcher Shijia Liao, Fish Audio has rapidly grown since launching last year. The company says more than 8 million people now use either the open-source or hosted versions of its AI voice models, while annual recurring revenue has reached $21 million.

Fish Audio focuses on building highly expressive and controllable AI-generated voices for a wide range of applications. Its platform includes a library of more than 15,000 natural language controls designed to meet the needs of creators, gaming studios and enterprises developing AI-powered customer interactions.

From an open-source project to a commercial platform

The company began as a personal project after Liao became dissatisfied with the limited expressiveness of synthetic voices available on the market. He trained an early voice-generation model using a single GPU before releasing it as open-source software.

That project evolved into the Fish Speech repository on GitHub, which has accumulated more than 31,000 stars and attracted developers, game creators and independent software builders.

During the past year, Fish Audio introduced five AI models, including four speech-generation models and one speech-to-text model. Three of its speech-generation models remain open source, while its latest S2.1 Pro model is available through the company's paid API service.

Serving creators and enterprise customers

Fish Audio offers subscriptions that provide creators and teams with voice-generation minutes and voice-cloning capabilities. It also supplies enterprise APIs and platform services that organisations including HeyGen and Sanas already use.

Chief executive and co-founder Rissa Cao said different industries require different voice characteristics. AI avatar providers prioritise realistic speech, gaming companies need expressive character voices, while businesses developing AI voice agents value low-latency responses and natural conversations.

Strengthening voice ownership protections

The startup has expanded its voice library partly by allowing users to contribute recordings for model training while compensating contributors when their voices are used commercially. However, the programme drew criticism after some creators alleged their voices had been uploaded without permission.

Cao said Fish Audio has since automated its takedown process. According to the company, creators can now submit a short voice sample or documentation proving ownership, allowing disputed voices to be removed from the platform in less than three minutes.

Even with the faster system, unauthorised uploads remain possible until a creator becomes aware of them and requests removal. Investors argue that long-term success for community-driven AI voice platforms will depend on stronger consent mechanisms, transparent licensing and revenue-sharing models that reward creators when their voices are used.

Funding to support the next generation of AI models

Cao said Fish Audio previously operated efficiently without outside investment while focusing on open-source development and creator tools. Growing demand from enterprise customers and the company’s ambition to build more advanced AI models ultimately led it to seek external funding.

The company plans to release an audio understanding model later this year and is also developing a speech-to-speech model as it broadens its product portfolio.

Fish Audio operates in an increasingly competitive market alongside companies including ElevenLabs, WellSaid, Cartesia, Speechify, Async and Krisp. Investors backing the startup believe its ability to train cost-efficient models and offer detailed developer controls will help it compete with larger AI companies while continuing to narrow the gap between synthetic and human-like speech.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav reports on startups, technology policy, and other significant technology-focused developments in India for TechAmerica.Ai. She previously worked as a research intern at ORF.