Mistral releases Voxtral TTS, its first open-weight speech model
Original titleOur first speech model, Voxtral TTS, is out. It delivers SOTA performance while significantly reducing cost compared to existing solution...
AISummary
Mistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.
Source: Guillaume Lample · x.comPublished · added here