Raschka reviews DeepSeek V4.1-Flash's encoder-decoder architecture overhaul
Original titleBig overhaul on DeepSeek V4.1 using an encoder-decoder setup.
AISummary
Sebastian Raschka says DeepSeek V4.1 contains a major architecture overhaul using an encoder-decoder setup, and he argues it could have been named V5.
The attached diagrams compare DeepSeek V4-Flash (284B) with DeepSeek V4.1-Flash (552B), which has 1M supported context and a 10-layer encoder.
The attached charts report a global KV cache per token of 890 bytes for V4.1-Flash, versus 3,514 for V4-Flash and 48,068 for DeepSeek-V3.2.
Source: Sebastian Raschka · x.comPublished · added here