Fish Audio just made S2.1 Pro free for a month.
Here's everything you need to know about it
- Clones any voice from 10 to 15 seconds of audio.
- ~90ms response, fast enough for real conversation
- 83 languages, one model.
- Word-level control over pronunciation and pauses
- Also has open weighted models
Under the hood it uses a Dual-AR design: a 4B-parameter model works out what to say and how it should feel, and a 400M-parameter model fills in the fine sound detail. That split is why it's both fast and expressive.
It's built from Fish Speech, their open-source project with tens of thousands of GitHub stars, and they still ship open-weight models you can self-host.
Price: roughly 1/6th of ElevenLabs for the same output.
The real pitch isn't a clean 10-second clip. It's a voice that survives a whole conversation, interruptions, corrections, laughter and language switches included.
🧵 1.