The V4-Pro preview went live early last Friday morning.
This morning, DeepSeek posted an announcement on its API dashboard: V4-Pro 75% off, valid until May 5. At the same time, the price for input cache hits across the entire API was slashed to one-tenth of the original.
From launch to steep discount, less than 72 hours passed.
Putting the Numbers on the Table
Original DeepSeek-V4-Pro pricing:
| Item | Original Price (per 1M tokens) |
|---|---|
| Input | $0.145 |
| Output | $3.48 |
| Cache hit | 0.025 yuan / million tokens (already reduced, 1/10 of original) |
During the promotion, V4-Pro's input price drops to roughly $0.036/M tokens — cheaper than anyone else. This price is so low that many suspect DeepSeek is running at a loss to drive user volume.
For comparison, current list prices of mainstream closed-source flagship models:
- GPT-5.5: Standard $1.5/M input, $10/M output
- Gemini 3.1 Pro: ~$1.25/M input, $10/M output
- Claude Opus 4.7: $15/M input, $75/M output
In other words, V4-Pro's original price already undercuts GPT-5.5 by more than ten times and Claude Opus 4.7 by more than a hundred times. Adding a 75% discount widens the gap further.
The Real Blow Is the Cache Cut
Looking only at V4-Pro's per-token price, it's already cheap. But the more aggressive move is DeepSeek slashing the input cache-hit price across the entire API product line to one-tenth of the original.
What cache-hit pricing is used for:
- Agent applications repeatedly running the same system prompt — large chunks are duplicated
- RAG scenarios frequently feeding the same context to the model
- Multi-turn conversations where history messages are repeatedly sent
In these scenarios, the cache-hit price is what enterprise users actually pay the bulk of. A 90% cut means the cost barrier for developers migrating from GPT and Claude to DeepSeek has been nearly flattened.
Why Strike at This Moment
After the V4-Pro preview launched early last Friday (April 24), the market reaction was mixed. Bloomberg's headline that day read "DeepSeek's Long-Awaited New Model Fails to Narrow US Lead in AI." Fortune reported that V4-Pro still lags behind Gemini 3.1 Pro on world knowledge benchmarks.
Simply put: technically not a game-changer, but price-wise still a contender.
DeepSeek's strategy is clear — since it can't directly challenge GPT-5.5 and Claude Opus at the very top in the short term, it takes another path: crush developers' migration costs to the point they can't ignore. Free model trials aren't new, but cutting prices by 75% within 72 hours of launch is a statement — telling every developer still on the fence: I'll save you the cost of switching.
Those familiar with the industry know DeepSeek used the same playbook when it launched R1 last January: sufficient capability, crushing prices. That wave triggered a price-cutting spree among Chinese open-source models, and OpenAI, Anthropic, and Google all adjusted prices multiple times afterward.
This time, it looks like the same script.
Who Will Crack First
Domestic competitors have already reacted. MiniMax's Hong Kong-listed shares fell 9% intraday, and Knowledge Atlas (Zhipu) dropped a similar amount. It's not because their models are bad — it's because their pricing room has been flattened by V4-Pro's promotional price.
In plain English: if you're still trying to sell your domestic model at $1/M tokens today, developers open DeepSeek's console and see $0.036 per million tokens — how do you explain being 27 times more expensive?
OpenAI and Anthropic likely won't follow in the short term. Their strategy has always been "high price, high quality" — customers pay extra for reliability, SLAs, compliance, and enterprise support. But the real pain is for the middle tier:
- Closed-source but mid-priced players like Mistral and Cohere
- Domestic players like Zhipu, MiniMax, and Moonshot
- A host of so-called "cost-effective" smaller vendors
They can neither claim to be cheaper than DeepSeek nor stronger than GPT-5.5. This squeeze blurs their positioning.
The Side Effect of This Move
DeepSeek's approach has a clear cost: margins may be negative.
V4-Pro is a 1.6-trillion-parameter MoE model, activating 49 billion parameters per inference. Running at this scale on Huawei Ascend 950PR hardware, the marginal cost per token won't be much higher than Gemini Flash or GPT-5.4-Nano — but compared to the promotional price, it's almost certainly a loss.
So why do it? The answer is likely that DeepSeek is treating "user volume" and "usage frequency" as core data points for its next funding round and potential IPO. A few days ago, reports emerged that Liang Wenfeng has begun reaching out to external investors — the first time he's opened the door after three years of refusing VC. Model quality has reached the second tier; the next story to tell is "developer market share."
Cheapness alone is not a moat. Cheap + good enough quality + open-source weights + stable API is. If V4-Pro's Hugging Face downloads and real API call volumes continue to surge over the next month, this move will have been worth it.
As for whether others will follow? We should know by June.
Sources: DeepSeek Slashes Fees for New AI Model in Chinese Price War (Bloomberg); DeepSeek slashes AI model costs, reignites price war in sector (Investing.com / Reuters); DeepSeek V4-Pro Price Cut: 75% (TheNextWeb); China's DeepSeek Slashes AI Prices With 75% Discount on V4-Pro Model (Republic World); CocoLoop, DeepSeek's Long-Awaited New Model Fails to Narrow US Lead in AI (Bloomberg)