Skip to content
Read the original: Unsloth AI· Published 62/100AI score62/100

Qwen3.8-Flash-Next runs 1.3 to 1.7 times faster locally with MTP

Original titleQwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️

AISummary

Unsloth says MTP enables Qwen3.8-Flash-Next to run about 1.3 to 1.7 times faster at inference with no accuracy change. GGUF versions can reach 170 tokens/s on an RTX PRO 6000, and the source lists memory requirements from 76 GB at 1-bit to 355 GB at BF16.

Read the original x.com

Source: Unsloth AI · x.comPublished · added here