Skip to content
Read the original: Georgi Gerganov· Published 29/100AI score29/100

Upgrade Qwen3.8-27B to DFlash for extra llama.cpp speed

Original titleIf you are using Qwen3.8-27B + MTP, make sure to upgrade to DFlash for extra speed:

AISummary

Georgi Gerganov says users of Qwen3.8-27B with MTP can get extra speed by switching to DFlash speculative decoding in llama.cpp. The command uses --spec-type draft-dflash with --spec-draft-n-max 7, and it requires the latest llama.cpp v0.6.0.

Read the original x.com

Source: Georgi Gerganov · x.comPublished · added here