Qwen3.8-Flash-Next runs at 68.3 tok/s on a single RTX 5090
AIA Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.