Read the original: FunAudioLLM (Alibaba Tongyi) · new models on Hugging Face·Published AI score42/100
Fun-CosyVoice3-0.5B-2512 Released as Open-Source Multilingual Text-to-Speech Model
Original titleFunAudioLLM/Fun-CosyVoice3-0.5B-2512
AISummary
Alibaba's FunAudioLLM has released Fun-CosyVoice3-0.5B-2512, a 0.5B-parameter LLM-based text-to-speech model on Hugging Face, with an RL variant also published.
The model supports zero-shot voice cloning across 9 languages and 18+ Chinese dialects and accents, with streaming output at latency as low as 150ms.
On the source's test-en benchmark, it reports a 2.24% WER and 71.8% speaker similarity, and the RL version reports 1.68% WER.
Source: FunAudioLLM (Alibaba Tongyi) · new models on Hugging Face · huggingface.co