🎙️ Build with VoxCPM2: NekoAudio - Let Your Catgirl AI Speak 🐱✨
What if your AI character could truly have a voice?
Developer @MindLiu67966 created NekoAudio, an open-source catgirl audio dataset designed for character voice modeling and TTS fine-tuning.
Built from NekoQA-30K, the dataset: 🐾 Removes stage directions/action descriptions from dialogue text for more natural speech 🎧 Includes 80K short audio samples (10-30s each) optimized for TTS fine-tuning 🗣️ Provides audio-text pairs for exploring character voice adaptation and expressive TTS
The dataset includes: 📦 Neko_Audio-80K_Short - ~100GB short audio dataset for efficient fine-tuning 🔗: https://huggingface.co/datasets/liumindmind/Neko_Audio-80K_Short 📦 Neko_Audio-30K_Long - ~200GB long-form audio dataset for extended speech generation 🔗: https://huggingface.co/datasets/liumindmind/Neko_Audio-30K_Long
The creator also fine-tuned VoxCPM2-2B with NekoAudio-80K and open-sourced the demo model NekoCPM2_demo.
A fun exploration of how open-source voice models can bring AI characters to life. 🚀
Powered by VoxCPM2: 🤗 Hugging Face: https://huggingface.openbmb.com/model/openbmb/VoxCPM2