Skip to content
Read the original: DeepSeek · new models on Hugging Face·Published PickAI score78/100

DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

Original titledeepseek-ai/DeepSeek-V4.1-Flash

AISummary

DeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

AIWhy it matters

The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

Read the original huggingface.co

Source: DeepSeek · new models on Hugging Face · huggingface.co