Rohan Paul@rohanpaul_ai
53AI 编辑部评分,满分 100
2026-07-30 22:53· 19分钟前
跳到正文
AI 摘要

Fish Audio 宣布 S2.1 Pro 免费使用一个月。该模型可从 10-15 秒音频克隆任意语音,响应约 90ms,支持 83 种语言,并提供词级发音与停顿控制。其 Dual-AR 架构包含 4B 参数模型(决定内容与情感)和 400M 参数模型(填充声音细节),价格约为 ElevenLabs 的 1/6。

Fish Audio just made S2.1 Pro free for a month.

Here's everything you need to know about it

  • Clones any voice from 10 to 15 seconds of audio.
  • ~90ms response, fast enough for real conversation
  • 83 languages, one model.
  • Word-level control over pronunciation and pauses
  • Also has open weighted models

Under the hood it uses a Dual-AR design: a 4B-parameter model works out what to say and how it should feel, and a 400M-parameter model fills in the fine sound detail. That split is why it's both fast and expressive.

It's built from Fish Speech, their open-source project with tens of thousands of GitHub stars, and they still ship open-weight models you can self-host.

Price: roughly 1/6th of ElevenLabs for the same output.

The real pitch isn't a clean 10-second clip. It's a voice that survives a whole conversation, interruptions, corrections, laughter and language switches included.

🧵 1.

Rohan Paul · @rohanpaul_ai · X·2026-07-30 22:53·19分钟前
在 X 看原推· x.com
AI 摘要

Fish Audio 宣布 S2.1 Pro 免费使用一个月。该模型可从 10-15 秒音频克隆任意语音,响应约 90ms,支持 83 种语言,并提供词级发音与停顿控制。其 Dual-AR 架构包含 4B 参数模型(决定内容与情感)和 400M 参数模型(填充声音细节),价格约为 ElevenLabs 的 1/6。

Fish Audio just made S2.1 Pro free for a month.

Here's everything you need to know about it

  • Clones any voice from 10 to 15 seconds of audio.
  • ~90ms response, fast enough for real conversation
  • 83 languages, one model.
  • Word-level control over pronunciation and pauses
  • Also has open weighted models

Under the hood it uses a Dual-AR design: a 4B-parameter model works out what to say and how it should feel, and a 400M-parameter model fills in the fine sound detail. That split is why it's both fast and expressive.

It's built from Fish Speech, their open-source project with tens of thousands of GitHub stars, and they still ship open-weight models you can self-host.

Price: roughly 1/6th of ElevenLabs for the same output.

The real pitch isn't a clean 10-second clip. It's a voice that survives a whole conversation, interruptions, corrections, laughter and language switches included.

🧵 1.

在 X 查看原推x.com