Skip to content
Read the original: Sebastian Raschka· Published 62/100AI score62/100

Raschka reviews DeepSeek V4.1-Flash's encoder-decoder architecture overhaul

Original titleBig overhaul on DeepSeek V4.1 using an encoder-decoder setup.

AISummary

Sebastian Raschka says DeepSeek V4.1 contains a major architecture overhaul using an encoder-decoder setup, and he argues it could have been named V5.

The attached diagrams compare DeepSeek V4-Flash (284B) with DeepSeek V4.1-Flash (552B), which has 1M supported context and a 10-layer encoder.

The attached charts report a global KV cache per token of 890 bytes for V4.1-Flash, versus 3,514 for V4-Flash and 48,068 for DeepSeek-V3.2.

Read the original x.com

Source: Sebastian Raschka · x.comPublished · added here