Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains
Original titleFull-Pipeline Inference Optimization for MiMo-V2.5 Series
Xiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention.
The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations.
It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.
The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.
Source: Xiaomi MiMo · mimo.xiaomi.com