Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model
AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.
Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.