Read the original: Ahead of AI (Sebastian Raschka)· Sebastian Raschka, PhD· Published · added 62/100AI score62/100
Recent LLM architecture changes that cut long-context KV cache and attention cost
Original titleRecent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
AISummary
Sebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs.
He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4.
The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.
Source: Ahead of AI (Sebastian Raschka) · magazine.sebastianraschka.comPublished · added here