Qwen3.8-Flash-Next runs 1.3 to 1.7 times faster locally with MTP
Original titleQwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️
AISummary
Unsloth says MTP enables Qwen3.8-Flash-Next to run about 1.3 to 1.7 times faster at inference with no accuracy change. GGUF versions can reach 170 tokens/s on an RTX PRO 6000, and the source lists memory requirements from 76 GB at 1-bit to 355 GB at BF16.
Source: Unsloth AI · x.comPublished · added here