SGLang now supports Kandinsky 6.0 Video from @kandinskylab_ai! It generates video and synchronized audio together from text or an image. Speech, sound, and lip-sync come out of a single model.
2 sizes: Lite (3B) and Pro (29B) with text-to-audio-video and image-to-audio-video modes
Built-in super-resolution up to 1920×1080
Dual-stream CrossDiT: a pretrained video stream and a new audio stream, linked by bidirectional cross-attention
Run it with SGLang Diffusion:

