Recompute SSM states instead of storing them for faster hybrid decoding
Original titleAs hybrid models (Qwen 3.5 / Nemotron Ultra) run agents with massive context, Gated-DeltaNet / Mamba states become a bottleneck. A simple...
AISummary
Tri Dao's post proposes loading SSM states, computing with them, and not storing them to make hybrid model decoding about 2x faster. The trick aims to ease the bottleneck from Gated-DeltaNet and Mamba states in hybrid models such as Qwen 3.5 and Nemotron Ultra, unlocking speculative decoding for SSMs.
Source: Tri Dao · x.comPublished · added here