DeepSeek V4 Flash makes capable AI agents cheaper to access through a hosted service. Its weights carry an MIT license, and its API costs $0.14 per million uncached input tokens and $0.28 per million output tokens. What the evidence does not establish is the viral claim that a comparable agent can run on a $5,000 machine.
DeepSeek supports local serving and publishes the model’s weights. Nvidia’s published serving recipe for V4 Flash targets datacenter hardware, with four H200 GPUs on a Kubernetes cluster for its aggregated deployment, rather than a single workstation. Neither DeepSeek’s model card and API documentation nor Nvidia’s deployment recipe shows the model running on a single consumer-priced machine, and no independent test of that setup has been published.
The clearer bargain is therefore DeepSeek’s hosted API. It offers a one-million-token context window, tool calling and Responses API support. DeepSeek says it plans to charge twice its regular prices during future peak hours, though it has not announced when that policy will begin.

Independent testing suggests the updated model is unusually capable for the portion used to process each token. Artificial Analysis gives DeepSeek V4 Flash 0731 an Intelligence Index score of 50, up from 40 for the April version.
DeepSeek’s earlier comparison with its larger V4 Pro showed Flash close on one software-engineering test but farther behind on a terminal-use test. The company’s later change log says the 0731 update surpassed the Pro preview on agent benchmarks, including a score of 82.7 on Terminal Bench 2.1, measured on DeepSeek’s own harness at maximum effort. Artificial Analysis separately reported 79% on that test and placed the update six points ahead of V4 Pro on its Intelligence Index.
Those results indicate that Flash can perform strongly on agent tasks, but they do not prove reliable everyday performance or modest hardware requirements. Ahmad Osman, founder of Osmantic AI, said in an interview relayed by MTS that V4 Flash “would easily run on a single DGX Spark that costs $5,000,” with performance “similar to Opus 4.6.” Artificial Analysis’s own leaderboard chart puts Opus 4.6 at 61 against 50 for V4 Flash 0731, an 11-point gap on the intelligence half of that comparison; the hardware half is undemonstrated.
Sources (5)
- Change Log | DeepSeek API Docs api-docs.deepseek.com
- DeepSeek V4 Preview Release | DeepSeek API Docs api-docs.deepseek.com
- Models & Pricing | DeepSeek API Docs api-docs.deepseek.com
- Deepseek-ai/DeepSeek-V4-Flash huggingface.co
- DeepSeek-V4-Flash | NVIDIA Dynamo Documentation docs.nvidia.com