How MoE inference splits into prefill, midfill, and decode regimes
Original titleComputation and Data Movement for Inference
AISummary
The article explains how Mixture of Experts models change inference by making prefill, midfill, decode attention, and decode experts distinct workloads. It describes how KV cache state, expert routing, and parallelism choices shape compute, memory, and network demands across an inference cluster.
Source: SemiAnalysis · newsletter.semianalysis.comPublished · added here