Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s
Original title (Chinese)
12GB 显存显卡跑 125B Qwen3.8 模型:Strata 登场,单张 RTX 5070 跑出 94 词元 / 秒
AISummary
Developer Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding.
On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.
Source: IThome · AI · ithome.comPublished · added here