Live video streaming needs to respond instantly while remaining visually consistent over long conversations.
Original titleLive video streaming needs to respond instantly while remaining visually consistent over long conversations.
AISummary
We achieved this by distilling a large 40-step diffusion teacher with 3-way CFG (120 evaluations per video chunk) into an unguided 2-step causal student with a fixed-length KV cache. Self-forcing helps the student resist drift and maintain near teacher quality with 60x fewer evaluations.
Source: AI at Meta · x.com