Meta unveils Muse Realtime Voice and Avatar with shared speech-token streaming
Original titleMuse Realtime Voice and Muse Realtime Avatar form a unified streaming architecture connecting conversational intelligence, voice, and vid...
AISummary
Meta's Muse Realtime Voice generates speech tokens encoding both content and prosody, and Muse Realtime Avatar consumes that shared stream to produce streaming video. Using a fixed-length history as motion context keeps computation bounded regardless of conversation length while synchronizing voice, lip motion, and expressions.
Source: AI at Meta · x.com