DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model
Original titledeepseek-ai/DeepSeek-V4-Flash-Vision-Exp
DeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules.
It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable.
The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.
The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.
Source: DeepSeek · new models on Hugging Face · huggingface.coPublished · added here