DeepSeek V3: 671B Parameters, Only 37B Activated

In December last year, DeepSeek released V3. What made competitors most uneasy wasn't the performance—it was the cost.

Architecture First

The core design of V3 is Mixture of Experts:

  • Total parameters: 671B
  • Actual activation per token: 37B
  • 256 experts per layer, 8 selected each time

It's like having 256 specialist doctors, but for each consultation you only call in the 8 most relevant ones. Large parameter count with manageable compute cost—that's the essence of MoE.

Benchmark Scores

  • MMLU: 87.1% (GPT-4o level)
  • MATH-500: 90.2%
  • Codeforces: 51.6th percentile
  • GPQA Diamond: 59.1%

These results place it among the top-tier open-source models as of late 2025.

$5.5 Million to Train a Top-Tier Model

V3's training cost was approximately $5.5M. In the same period, U.S. companies spent hundreds of millions to train models of similar scale. When this number came out, the entire industry started rethinking: is DeepSeek just extremely frugal, or has everyone else been wasting money?

Cost-saving secrets include:

  • FP8 mixed-precision training, cutting memory bandwidth requirements in half
  • Strategic use of spot instances, reducing compute costs by 70%
  • The MoE architecture itself is a natural compute saver

Why It Matters

One month after V3's release, DeepSeek dropped R1. Together, these two models essentially proved one thing: frontier AI research doesn't necessarily require burning billions of dollars. For the global AI competition landscape, this signal is more important than any single benchmark.

Sources: CocoLoop, DeepSeek V3 technical report, BentoML technical guide