llama.cpp adds DFlash speculative decoding support for faster local inference
Original titlellama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, t...
AISummary
llama.cpp has added DFlash support to its speculative decoding options, joining MTP, Eagle3, and various ngram-based techniques. The author says these methods further improve local model performance. The change was developed with thanks to the NVIDIA team and Ruixiang Wang.
Source: Georgi Gerganov · x.comPublished · added here