In December last year, DeepSeek released V3. What made competitors most uneasy wasn't the performance—it was the cost.
Architecture First
The core design of V3 is Mixture of Experts:
- Total parameters: 671B
- Actual activation per token: 37B
- 256 experts per layer, 8 selected each time
It's like having 256 specialist doctors, but for each consultation you only call in the 8 most relevant ones. Large parameter count with manageable compute cost—that's the essence of MoE.
Benchmark Scores
- MMLU: 87.1% (GPT-4o level)
- MATH-500: 90.2%
- Codeforces: 51.6th percentile
- GPQA Diamond: 59.1%
These results place it among the top-tier open-source models as of late 2025.
$5.5 Million to Train a Top-Tier Model
V3's training cost was approximately $5.5M. In the same period, U.S. companies spent hundreds of millions to train models of similar scale. When this number came out, the entire industry started rethinking: is DeepSeek just extremely frugal, or has everyone else been wasting money?
Cost-saving secrets include:
- FP8 mixed-precision training, cutting memory bandwidth requirements in half
- Strategic use of spot instances, reducing compute costs by 70%
- The MoE architecture itself is a natural compute saver
Why It Matters
One month after V3's release, DeepSeek dropped R1. Together, these two models essentially proved one thing: frontier AI research doesn't necessarily require burning billions of dollars. For the global AI competition landscape, this signal is more important than any single benchmark.
Sources: CocoLoop, DeepSeek V3 technical report, BentoML technical guide