Skip to content
View original post on X: SGLang· 62/100AI score62/100

SGLang adds support for Kandinsky 6.0 Video audio-visual generation

AISummary

SGLang now supports Kandinsky 6.0 Video, which generates video and synchronized audio together from text or an image. The model comes in Lite (3B) and Pro (29B) sizes, with built-in super-resolution up to 1920×1080. A sample sglang serve command for the Pro distilled model is included.

Post on XView on X
@sgl_project

SGLang now supports Kandinsky 6.0 Video from @kandinskylab_ai! It generates video and synchronized audio together from text or an image. Speech, sound, and lip-sync come out of a single model.

2 sizes: Lite (3B) and Pro (29B) with text-to-audio-video and image-to-audio-video modes
Built-in super-resolution up to 1920×1080
Dual-stream CrossDiT: a pretrained video stream and a new audio stream, linked by bidirectional cross-attention

Run it with SGLang Diffusion:

KandinskyLab.AI@kandinskylab_ai
🚀 Kandinsky 6.0 Video is open source! Meet Lite (3B) and Pro (29B) for video generation with synchronized audio. Code and weights are released under MIT. Try them out and share what you create! https://github.com/kandinskylab/kandinsky-6
View quoted post on X

Source: SGLang · x.comPublished · added here