Tri Dao Shares Speculative Speculative Decoding, a Claimed Up-to-2x LLM Inference Speedup
Original titleAttack of the asynchronous machines. We’ve seen this a lot in GPU kernels. This time the same principle applies in speculative decoding
AISummary
Tri Dao reposts a quoted post from @tanishqkumar07 introducing Speculative Speculative Decoding (SSD), an LLM inference algorithm claimed to be up to 2x faster than leading inference engines.
The quoted post credits collaborators @tri_dao and @avnermay and links to a thread with details. Tri Dao's own text says the approach applies an asynchronous-machines principle seen in GPU kernels to speculative decoding.
Source: Tri Dao · x.comPublished · added here