How GLM-5.3 Sparse Attention Affects HBM and Serving Costs on GB200, GB300, and MI355X
Original titleHow GLM5.3 Sparse Attention Affects HBM Memory Usage
AISummary
Sparse attention cuts per-operation KV cache reads but does not reduce overall memory capacity, so top-k cache misses still depend on HBM. SemiAnalysis's InferenceX estimates GB200 at about $0.044 per million total tokens at 150 tokens per second, roughly 12% below MI355X running ATOM at $0.049.
Neither system holds a uniform cost advantage across the tested 100, 125, and 150 tokens-per-second targets.
Source: SemiAnalysis · newsletter.semianalysis.comPublished · added here