Palo Alto-based Fish Audio has secured $50 million in a seed funding round to further develop its advanced AI voice models. The company aims to provide increasingly expressive and steerable synthetic voices for both creative applications and enterprise solutions.
Advancing AI Voice Generation Capabilities
The demand for artificial intelligence-generated voice models is substantial, driven by diverse needs across sectors. For creative endeavours, these models require enhanced expressiveness to imbue content with emotion and nuance. Concurrently, enterprises seeking to automate customer support and sales operations require voices that are highly steerable and adaptable to specific scenarios.
Fish Audio intends to address these varied requirements with its extensive library of over 15,000 natural language controls. Since its inception last year, the startup has seen remarkable adoption, with more than 8 million individuals utilising its open-source or hosted models. The company has also achieved significant commercial traction, generating $21 million in annual recurring revenue.
Seed Funding Fuels Future Development
To sustain its rapid growth and further innovation, Fish Audio announced on Tuesday the closure of its $50 million seed funding round. The investment was led by Coreline Ventures and Capital Today, with additional participation from a notable group of venture firms including 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
The company's foundation was laid by former NVIDIA researcher Shijia Liao, who, motivated by the limitations of existing synthetic voices, developed and open-sourced an initial voice generation model using a single GPU. This project, the Fish Speech repository on GitHub, has garnered considerable attention, accumulating over 31,000 stars and finding widespread use among indie developers, video game designers, and content creators.
Expanding Product Offerings and Enterprise Adoption
Over the past year, Fish Audio has introduced five distinct models, comprising four for speech generation and one for speech-to-text conversion. While three of its speech generation models remain open-source, the company's latest offering, the S2.1 Pro model, is exclusively accessible through its paid API.
Fish Audio provides tiered monthly subscription plans designed for creators and teams, which include a set allocation of generation minutes and advanced voice cloning features. Additionally, the company offers an enterprise version of its APIs and platform. Organisations such as HeyGen, Sanas, and Plaud are already leveraging these solutions to enhance their operations.
Addressing Creator Concerns and Future Innovations
Fish Audio has historically built its voice library partly by inviting users to contribute their own voices for model training, offering compensation for their use. However, this approach led to some challenges, with certain creators alleging their voices were used without explicit consent a few months ago. While the company had a DMCA content take-down process, the resolution of these requests was often protracted.
In response to these concerns, Fish Audio's CEO and co-founder, Rissa Cao, stated that the take-down process has now been automated. Creators can swiftly prove ownership of their voice data through a short sample or contract, enabling voice removal from the platform in under three minutes. This move aims to bolster trust within its community-driven model, a sentiment echoed by investors.
Oskue Honda, a partner at Coreline Ventures, emphasised the critical importance of trust in a community-centric model, stating, "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts." He further suggested the industry should move towards verified voice ownership, clear licensing, and equitable revenue-sharing models.
Looking ahead, Fish Audio plans to launch an audio understanding model and a speech-to-speech model this year, signalling its commitment to expanding the capabilities of its AI voice technology. The company operates within a competitive market, with established players such as ElevenLabs, WellSaid, and Speechify also vying for market share.