Stability AI's SemanTok makes video world models more efficient
Original titleWhat if making video-based world models smaller isn't just about better compression, but about making their representations easier to pre...
AISummary
Stability AI's Interactive Research team introduced SemanTok, which makes early video tokens more semantically meaningful so the representation is easier to predict. According to the post, a model using SemanTok matches or beats the performance of a model more than three times its size. The approach targets more efficient autoregressive video generation.
Source: Stability AI · x.comPublished · added here