AI Price War: DeepSeek Drops V4-Flash, OpenAI Slashes 80%

The AI model market just entered a new phase. The price war is no longer a background trend. It is the story.
This week, two events landed almost simultaneously. DeepSeek released the public beta of its V4-Flash API, a model explicitly tuned for agent tasks. OpenAI responded by cutting GPT-5.6 Luna prices by 80 percent. The timing is not a coincidence. It is the clearest signal yet that the AI industry is being reshaped by a price war that is accelerating, not slowing down.
DeepSeek V4-Flash: Built for agents
DeepSeek's V4-Flash-0731 model is the public beta version of its lightweight agent-focused API. The model scored 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, benchmarks that measure agent task completion and software engineering capability. It adds support for the Responses API and is adapted for Codex, making it immediately usable for developers building AI-powered coding tools.
The model keeps the same architecture and size as the preview version released earlier, but has been retrained. The update applies only to the V4-Flash API; the V4-Pro API and the consumer-facing models on DeepSeek's app and website remain unchanged.
The focus on agent tasks is strategic. While DeepSeek's flagship V4-Pro (1.6 trillion parameters, Mixture-of-Experts) competes at the frontier, V4-Flash targets the fast-growing market for AI agents that browse the web, write code, and execute multi-step tasks autonomously. It is the same market OpenAI is chasing with Codex, and Anthropic is chasing with Claude Code.
OpenAI goes into China pricing mode
On the same day, OpenAI announced that GPT-5.6 Luna, its smallest and most affordable model, would drop to $0.20 per million input tokens and $1.20 per million output tokens, an 80 percent cut. Terra, the mid-tier model, fell 20 percent to $2 and $12. Sol pricing stayed the same.
OpenAI attributed the cuts to efficiency gains from GPT-5.6 Sol, its flagship model, which self-optimized GPU software and cut deployment costs by 20 percent while improving token generation by 15 percent through speculative decoding. But the subtext is clear: Chinese competition is forcing the pace.
Microsoft is now openly promoting its own MAI models as cheaper alternatives to OpenAI. The price pressure from DeepSeek, which made its V4-Pro pricing permanently 11 times cheaper than GPT-5.5, has not let up. DeepSeek is now the fastest-growing AI provider on Ramp, a US financial services firm that tracks spending across 50,000 companies.
The economics of a price war
The numbers tell a stark story. A task that cost one dollar on leading models a year ago now runs for about six cents on Luna, nearly nine times faster. DeepSeek V4-Pro costs roughly eleven times less than GPT-5.5 on input. The gap between frontier and budget models is shrinking, and the entire market is being pulled down the cost curve.
This is good for developers and startups. But it creates a structural problem for the AI labs themselves. They are locked in an infrastructure arms race. DeepSeek is raising a new round at a $71 billion valuation to build its own data centers and chips, on top of the $7 billion it raised in May. OpenAI, Anthropic, and Google are spending similar amounts. If revenue growth slows because prices keep falling, the balance sheets behind those investments start to look fragile.
The three-way war between DeepSeek, OpenAI, and Chinese rivals like Moonshot and Alibaba is not a temporary pricing promotion. It is a structural shift. The winners will be the labs that can deliver capable models at the lowest cost, not the ones with the most impressive benchmark scores.
What this means
The AI industry is entering a phase where cost efficiency matters as much as raw intelligence. DeepSeek's V4-Flash public beta is a bet on agent workloads at scale. OpenAI's 80 percent price cut is a defensive move. And the broader market is watching to see which labs can sustain the margins to keep investing in the next generation of models.
For developers, the calculus is simple: the best model for any given task keeps getting cheaper. For investors, the question is harder. When everyone is racing to the bottom on price, someone has to be able to afford the top.
Sources
- TechNode: DeepSeek puts V4-Flash API into public beta (https://technode.com/2026/07/31/deepseek-puts-v4-flash-api-into-public-beta/)
- The Decoder: OpenAI cuts GPT-5.6 Luna prices by 80 percent (https://the-decoder.com/openai-gpt-5-6-luna-price-cut-80-percent/)
- The Decoder: DeepSeek needs more cash just weeks after closing its first $7 billion round (https://the-decoder.com/deepseek-needs-more-cash-just-weeks-after-closing-its-first-7-billion-round/)
- Bloomberg: DeepSeek Unveils Public Beta API for Flagship AI Model (https://www.bloomberg.com/news/articles/2026-07-31/deepseek-unveils-public-beta-api-for-flagship-ai-model)
- VentureBeat: AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% (https://venturebeat.com/2026/07/30/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost/)
- Ramp (via Decoder): DeepSeek fastest-growing AI provider on Ramp spending data