Qwen3.8-Flash-Next runs at 68.3 tok/s on a single RTX 5090
Original titleStronger open models + new inference systems also means super powerful local AI is coming! Cool work by Shuo from UC Berkeley Sky Lab.
AISummary
A Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.
Source: Matei Zaharia · x.comPublished · added here