RoboTTT scales robot policy context to 8,000 timesteps with constant inference cost
Original titleWe scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot pol...
AISummary
Jim Fan introduced RoboTTT, a robot model that uses test-time training to compress history into a tiny inner model updated at each sensor reading.
The post reports closed-loop performance rising steadily from 128 to 8K timesteps, and 8K-context pretraining beating 1K by 62%.
It also claims one-shot in-context learning from human video and mid-episode error recovery, with learning continuing after deployment.
Source: Jim Fan · x.comPublished · added here