Skip to content

Fun-CosyVoice3-0.5B-2512 Released as Open-Source Multilingual Text-to-Speech Model

Original titleFunAudioLLM/Fun-CosyVoice3-0.5B-2512

AISummary

Alibaba's FunAudioLLM has released Fun-CosyVoice3-0.5B-2512, a 0.5B-parameter LLM-based text-to-speech model on Hugging Face, with an RL variant also published.

The model supports zero-shot voice cloning across 9 languages and 18+ Chinese dialects and accents, with streaming output at latency as low as 150ms.

On the source's test-en benchmark, it reports a 2.24% WER and 71.8% speaker similarity, and the RL version reports 1.68% WER.

Read the original huggingface.co

Source: FunAudioLLM (Alibaba Tongyi) · new models on Hugging Face · huggingface.co