Skip to content
View original post on X: Unsloth AI· Pick78/100AI score78/100

Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

AISummary

Unsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

AIWhy it matters

The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

Post on XView on X
@UnslothAI

Qwen3.8-Flash can now be run locally! 🔥

The 125B MoE model outperforms Claude-Opus-4.6 (Max).

Run on 75GB RAM via Unsloth GGUFs.

Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds.

Guide: https://unsloth.ai/docs/models/qwen3.8-next
GGUF: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF

Qwen@Alibaba_Qwen
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Blog: https://qwen.ai/blog?id=qwen3.8-flash-next - Technical Report: https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf - Hugging Face: https://huggingface.co/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921FsyOPe&file=Qwen3.8-Flash-Next - ModelScope: https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921XAP2dV&file=Qwen3.8-Flash-Next
View quoted post on X

Source: Unsloth AI · x.comPublished · added here