Skip to content
Read the original: Georgi Gerganov· 13/100AI score13/100

Gerganov shows llama serve running Qwen3.8-27B GGUF with MTP drafting

Original titlesimple:

AISummary

Georgi Gerganov shared a single command that serves the ggml-org/Qwen3.8-27B-GGUF model using llama serve with --spec-type draft-mtp for speculative decoding. The post is a brief command example and gives no benchmark, speed, or setup details.

Read the original x.com

Source: Georgi Gerganov · x.comPublished · added here