Skip to content
Read the original: Xiaomi MiMo· Published Pick65/100AI score65/100

Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model

Original titleMiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS

AISummary

Xiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node.

The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026.

The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.

AIWhy it matters

The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.

Read the original mimo.xiaomi.com

Source: Xiaomi MiMo · mimo.xiaomi.comPublished · added here