Back to stories
fundingGenerated by an AI editor from the reporting and web sources listed on this page.

Fish Audio lands $50M seed to make AI voice the default interface for every application

The Palo Alto startup, born from a former NVIDIA researcher's weekend project, already counts 8 million users and $21M ARR as it builds expressive, steerable voice models for creators and enterprises alike.

Published The total reporting and web sources attached to this story.How many attached sources came from wider web research rather than monitored news feeds.The AI editor’s assessment of how strongly the attached sources’ quality and agreement support this article.

What matters

  • Fish Audio raised a $50M seed round led by Coreline Ventures and Capital Today, with participation from at least eight other investors.
  • The startup has 8 million users and $21M in annual recurring revenue since launching commercially last year.
  • Its open-source Fish Speech repository has over 31,000 GitHub stars, built from founder Shijia Liao's initial single-GPU voice models.
  • The platform offers 15,000+ natural language controls targeting both creative expressiveness and enterprise steerability.
  • SiliconANGLE reported the round as $52M while TechCrunch reported $50M; the exact figure remains slightly unclear.

Funding facts

Amount:
$50 million (TechCrunch) / $52 million (SiliconANGLE)
Round:
Seed
Lead investors:
Coreline Ventures, Capital Today

Open source

Repository:
https://github.com/fishaudio/fish-speech

What happened

Fish Audio, a Palo Alto–based AI voice startup officially incorporated as Hanabi AI Inc., announced on July 28, 2026 that it has raised $50 million in a seed round led by Coreline Ventures and Capital Today. The round also drew participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, HF0, and 645 Ventures, among other backers. (SiliconANGLE reported the figure as $52 million; TechCrunch reported $50 million.)

The company was founded by Shijia Liao, a former NVIDIA video researcher who grew frustrated with the flat, monotonous synthetic voices produced by early AI models. Liao began training his own text-to-speech and voice cloning models on a single GPU in his bedroom, eventually open-sourcing the work as Fish Speech. That repository has since accumulated more than 31,000 stars on GitHub.

Since launching commercially last year, Fish Audio has attracted more than 8 million users across its open-source and hosted offerings and now generates $21 million in annual recurring revenue. The platform offers a library of more than 15,000 natural language controls, allowing users to fine-tune voice output for both creative expressiveness and enterprise-grade steerability.

Why it matters

AI-generated voice is becoming a critical interface layer across consumer and enterprise applications. Creators need voices that convey emotion and nuance, while enterprises automating customer support and sales operations need voices that are predictable and controllable. Fish Audio is betting that a single platform with deep natural language controls can serve both audiences.

The startup's traction — 8 million users and $21 million ARR within roughly a year of launch — suggests strong product-market fit in a category that has attracted heavy competition. The sizable seed round, backed by a broad syndicate of investors, signals that venture capital sees meaningful upside in voice as a default interface for AI models, not just a niche tool for content creators.

Fish Audio's open-source roots also matter. The Fish Speech repository's 31,000+ GitHub stars indicate a developer community that could accelerate adoption, contribute improvements, and create a moat against proprietary competitors.

What to watch

  • How Fish Audio deploys the $50 million: the company has not yet detailed specific hiring, infrastructure, or product roadmap plans tied to the raise.
  • Whether the platform can maintain its open-source momentum while scaling hosted enterprise offerings — a balance that has tripped up other open-source-first startups.
  • Competitive dynamics in the AI voice space, where larger players and well-funded startups are all chasing creators and enterprises simultaneously.
  • The discrepancy in reported round size ($50M vs. $52M) between TechCrunch and SiliconANGLE — a minor but notable detail that may reflect undisclosed additional capital or a reporting lag.

What to do next

Developers

Explore the Fish Speech open-source repository on GitHub and evaluate its text-to-speech and voice cloning capabilities for your projects.

The repo has 31,000+ stars and offers a community-tested foundation for voice AI integration.

Founders

Study Fish Audio's open-source-to-commercial path as a template for building developer traction before monetizing hosted services.

The startup reached $21M ARR within roughly a year by leveraging open-source adoption as a funnel into paid offerings.

PMs

Assess Fish Audio's 15,000+ natural language controls as a benchmark for how granular voice customization should be in AI-powered products.

The control library targets both creative and enterprise use cases, offering a model for product differentiation in voice AI.

Investors

Note the broad syndicate and strong ARR traction, but scrutinize how the company plans to defend against larger, well-capitalized voice AI competitors.

$21M ARR in year one is impressive, but the AI voice market is crowded and competitive dynamics will determine long-term value.

Operators

Pilot Fish Audio's hosted models for customer support or sales automation workflows where expressive, steerable voice output matters.

The platform's enterprise focus on controllability could reduce the uncanny-valley problem in automated voice interactions.

How to test

  1. 1Clone the Fish Speech repository from GitHub and review the README for setup instructions.
  2. 2Install dependencies and download a pre-trained model checkpoint as documented.
  3. 3Run a basic text-to-speech inference on sample text to evaluate voice quality and expressiveness.
  4. 4Experiment with natural language controls from the 15,000+ control library to adjust tone, emotion, and pacing.
  5. 5Compare output quality against at least one competing open-source or commercial voice model.

Caveats

  • The open-source model may differ from Fish Audio's hosted enterprise version in quality and feature set.
  • Voice cloning capabilities raise ethical and legal considerations around consent and misuse; review the project's usage guidelines.
  • Reported round size differs between sources ($50M vs. $52M), which may indicate undisclosed details about the funding.