Gerganov shows llama serve running Qwen3.8-27B GGUF with MTP drafting
Original titlesimple:
AISummary
Georgi Gerganov shared a single command that serves the ggml-org/Qwen3.8-27B-GGUF model using llama serve with --spec-type draft-mtp for speculative decoding. The post is a brief command example and gives no benchmark, speed, or setup details.
Source: Georgi Gerganov · x.comPublished · added here