The Rundown AI homepage
Artificial intelligence/News & analysis

DeepSeek V4.1-Flash puts pressure on AI pricing

DeepSeek V4.1-Flash pairs low API prices with selected benchmark wins, giving developers a cheaper option while U.S. models retain the overall lead.

By The Rundown Editorial TeamReviewed by Kelly Pitts3 min read
DeepSeek turns up the pressure on AI pricing — newsletter story image
Image source: DeepSeek

DeepSeek released V4.1-Flash on September 10, pairing downloadable model weights with low API prices and selected benchmark wins in its own testing. As The Rundown reported, the launch adds pressure on AI pricing while U.S. labs retain the overall capability lead.

Prices and access

DeepSeek’s rate card lists $0.15 per million uncached input tokens and $0.60 per million output tokens outside peak hours. During peak hours, those rates double to $0.30 and $1.20. Peak windows run from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

Existing Flash API names already route to V4.1-Flash, DeepSeek says. V4-Pro requests are scheduled to move to Flash on September 14 at 04:00 UTC and will carry Flash rates until V4.1-Pro launches. The company has not announced a date for that model.

Developers can also download the weights under an MIT license, giving them the option to run Flash on their own infrastructure.

Benchmark wins with limits

DeepSeek’s results at maximum reasoning effort give Flash 74.2 on the DeepSWE v1.1 coding benchmark, ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0. The company also reports wins over both on AutomationBench and over GPT-5.6 Sol on CyberGym.

The same results put Flash behind both U.S. models on Terminal-Bench 4.0. It also trails V4-Pro on GPQA Diamond, scoring 90.9 against 92.4.

On the broader Artificial Analysis Intelligence Index, Flash scores 40 at maximum reasoning effort. The evaluator’s September 9 comparison placed GPT-6 Astra and Claude Fable 5.1 jointly first at 53.

Why it matters

DeepSeek’s efficiency and pricing remain its edge. Flash gives developers a reason to test cheaper models for coding and automated workflows even while U.S. labs lead on overall capability. Teams could move suitable tasks to Flash and keep paying for stronger models where the extra capability improves results. That leaves the competition for frequent, paid usage open.

DeepSeek reports that Flash’s global cache takes up roughly a quarter as much memory as V4-Flash’s. Lower memory needs could give it room to keep prices low. Developers who run the downloaded weights themselves would also need to factor in infrastructure and upkeep when comparing costs.

The practical test is the cost of completed work. Teams should measure how often a workflow succeeds, how many tokens it consumes, and what retries and tools add to the bill. Scheduling flexible jobs outside peak hours halves the listed token rates. Long reasoning outputs or repeated attempts could absorb some of that saving.

Chinese labs have already shown they can win usage. OpenRouter’s December 2025 study with a16z found that Chinese models with open weights reached nearly 30% of token usage on its platform in some weeks. The study covered November 2024 through November 2025. It also found that premium models retained substantial demand while cheaper models attracted buyers focused on cost. Whether buyers have become more sensitive to price since then remains unclear.

Distillation remains a possible factor in this competition. In its September threat report, Anthropic alleges that DeepSeek extracted Opus reasoning traces through methods designed to bypass its controls. Anthropic says its research shows distillation, or training on another model’s outputs, can improve coding, reasoning, and automated task performance. The report does not establish whether that activity contributed to V4.1-Flash or its prices.

Rivals could face pressure to cut rates or justify their premiums with better results. Whether that pressure lowers margins depends on which workloads move, what they cost to serve, and how much customers will pay for premium performance.

Sources & further reading

This story builds on reporting from The Rundown newsletter on September 11, 2026.