DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture
Original titleDeepSeek-V4.1-Flash: more efficient prefill for coding agents
AISummary
DeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input.
Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.
AIWhy it matters
The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.
Source: Baseten Blog · baseten.coPublished · added here