Mistral's Voxtral Realtime streams speech with sub-200ms latency and open weights
Original titleVoxtral Realtime is built for voice agents and live applications. Its natively streaming architecture delivers latency configurable to su...
AISummary
Voxtral Realtime is a natively streaming speech model for voice agents and live applications, with latency configurable down to sub-200ms. At 480ms it stays within 1-2% WER of the offline model, and the weights are released under Apache 2.0. The attached FLEURS chart compares word error rates across latency settings for ten languages, including Chinese.
Source: Guillaume Lample · x.comPublished · added here