DeepSeek V4 Flash Vision Exp is an experimental multimodal model built on the DeepSeek V4 Flash backbone, adding image understanding while fully matching the base model on text capabilities including agents, reasoning, and world knowledge. It uses a 284B-parameter Mixture-of-Experts architecture (13B active parameters per token) with a highly efficient vision encoder that compresses image representations into a fraction of the KV-cache footprint of competing models. On multimodal agent benchmarks, it makes a major leap over V4 Flash, bringing multimodal agent performance close to frontier-level models. It supports mixed text and image input via base64, external URLs, or the Files API, with a 1-million-token context window.
Evaluations
12
across 12 benchmarks
Latency
329ms
Context length
1M
tokens
Cost
$0.22 · $0.66
input · output per 1M tokens
Key takeaways
Key takeaways will appear here once evaluations are run.