DeepSeek-V4-Flash undercuts AI pricing: $0.28 per million tokens for agentic coding near GPT-level
This digest was compiled by AI from multiple sources — links to the originals are below.

DeepSeek's V4-Flash model achieves near-frontier performance on agentic coding benchmarks while costing just $0.28 per million output tokens, compared to $25 for Claude Opus 4.8. The model uses a Mixture-of-Experts architecture with hybrid sparse attention to slash serving costs without sacrificing quality.
Efficiency Architecture and Pricing
DeepSeek-V4-Flash leverages a Mixture-of-Experts design where only a fraction of parameters activate per token, and employs a hybrid sparse attention mechanism — combining compressed sparse attention and head-wise compressed attention — to handle 1 million-token contexts efficiently. This allows inference at a fraction of the cost of dense models: $0.28 per million output tokens versus $25 on Claude Opus 4.8 and $3–15 on GPT-4 competitors. The model’s input token cost is $0.14, making long-context tasks economically viable. DeepSeek’s approach demonstrates that optimization can narrow the performance gap with frontier models by 30–50% while reducing cost by 99%.
Agentic Coding and Practical Use
On SWE-bench Verified, a key agentic coding benchmark, Flash scores 36.6, placing it ahead of GPT-4.1-mini (31.7) and within 4 points of Claude Opus 4.8 (40.5). The model retains reasoning traces across tool-calling turns, enabling more coherent agent behaviors. It supports multiple reasoning-effort modes and an OpenAI-compatible API, easing migration from existing systems. However, Flash’s broad world knowledge lags behind larger models, making it less suitable for knowledge-heavy tasks. Operations teams must still sandbox bash tool calls and monitor for silent model updates.
What's Next
DeepSeek is expected to expand V4-Flash’s availability in regional data centers to address latency and data residency concerns. Whether its pricing pressures trigger a broader race to the bottom among model providers remains an open question, as competitors may prioritize premium features over cost.
1 source
DeepSeek-V4-Flash undercuts AI pricing: $0.28 per million tokens for agentic coding near GPT-level






