DeepSeek turned the model on through its developer service on Aug. 21. The model-access service OpenRouter added it the same day. It is a separate choice called DeepSeek V4 Flash Vision Exp, not an automatic upgrade to the text-only V4 Flash.
The model takes an image and a written prompt in one request, letting an AI tool read a screenshot, examine a chart or describe a picture directly. That removes an extra handoff used by at least one community project, which first had a separate vision model translate images into words for V4 Flash.
DeepSeek says Vision Exp keeps V4 Flash’s text abilities and approaches Anthropic’s Opus 4.8 on tests of agents working with images. OpenRouter says it wins on Agents’ Last Exam and ZeroBench, though DeepSeek’s own table has Opus 4.8 ahead on most of the remaining tests. Those launch scores have not been independently reproduced under matching settings, and the comparison is with an older generation: Anthropic released Opus 5 in July.
DeepSeek charges images as input tokens at V4 Flash’s existing rates: 22 to 44 cents per million input tokens and 66 cents to $1.32 per million output tokens, depending on the time of day. Because one image is capped at 384 input tokens, even the higher rate puts the image itself below two-hundredths of a cent before the model’s answer.