DeepSeek has put an upgraded V4-Flash API into public beta, claiming much stronger coding-agent performance while charging $0.14 per million uncached input tokens and $0.28 per million output tokens. Its pricing page says it will adopt peak-hour pricing at twice those rates, with the start date not yet announced.

Via @ArtificialAnlys

Artificial Analysis found the new build used about 12% fewer output tokens than its predecessor to run the same evaluation suite, which would cut the cost of repeated agent runs. It separately found a broad improvement, scoring the model 50 on its Intelligence Index, up 10 points, and placing it on its frontier for intelligence versus cost per task.

This is not Flash’s debut. The news is that the re-post-trained V4-Flash-0731 build entered public beta on July 31, 2026. DeepSeek says it changed neither the architecture nor the size of the preview model.

A smaller model takes a large step

Flash activates 13 billion parameters per token, compared with 49 billion for DeepSeek’s larger V4-Pro.

Despite that gap, DeepSeek says the build’s scores, 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, far exceed those of its larger V4-Pro-Preview. Artificial Analysis separately measured 79% on Terminal-Bench 2.1.

Via @ns123abc

DeepSeek ran its public coding tests at maximum effort with its own unreleased harness. Two other tests in its table are internal. The results are therefore an early signal rather than a settled independent comparison.

The practical advantage is agent access

Long-running agents repeatedly read context, call tools and produce code, accumulating costs at every step. Cached input costs $0.0028 per million tokens, a roughly 98% discount that Artificial Analysis called a key driver of Flash’s cost per task running about 60% below GPT-5.6 Luna (max), even after OpenAI’s price cut.

Flash also supports the Responses API and Codex integration. DeepSeek says V4-Pro supports neither yet, though both are expected in early August 2026. That gives Flash a temporary practical advantage beyond its benchmark scores.

The rollout is limited to the V4-Flash API. DeepSeek’s V4-Pro API, app and website models are unchanged.

The plain takeaway: retraining made DeepSeek’s cheap model substantially more capable for coding agents without making it larger.

Sources (12)