DeepSeek V4 Flash Vision Exp is an experimental multimodal model built on the DeepSeek V4 Flash backbone, adding image understanding while fully matching the base model on text capabilities including agents, reasoning, and world knowledge. It uses a 284B-parameter Mixture-of-Experts architecture (13B active parameters per token) with a highly efficient vision encoder that compresses image representations into a fraction of the KV-cache footprint of competing models. On multimodal agent benchmarks, it makes a major leap over V4 Flash, bringing multimodal agent performance close to frontier-level models. It supports mixed text and image input via base64, external URLs, or the Files API, with a 1-million-token context window.