OpenAI Bets $20 Billion on Cerebras

Two $20 billion chip deals emerged in the same week.

One was Nvidia's acquisition of Groq late last year for $20 billion — a company specializing in AI inference chips and Cerebras's main architectural rival.

The other is OpenAI's renewed commitment to Cerebras: an additional $20 billion in compute spending over three years, plus warrants for up to 10% equity, and a separate $1 billion dedicated fund to build data centers for Cerebras.

These two deals happening simultaneously is no coincidence. The AI computing battlefield is shifting from training to inference, and everyone is jockeying for position.

Inference Becomes the Main Battlefield

First, some context.

AI compute comes in two forms: training (feeding large amounts of data to a model to adjust parameters, a one-time cost) and inference (using a trained model to answer questions, an ongoing cost).

A few years ago, training was the primary expense. That's no longer the case.

In 2025, inference already accounts for 50% of total AI compute spending. By 2026, that share is expected to jump to two-thirds.

Every time you send a message to ChatGPT, Claude, or Gemini, it triggers an inference request. The more users, the higher the inference cost, and it scales linearly. Large model companies serving hundreds of millions of users need inference infrastructure efficiency to determine whether their gross margins are sustainable.

Nvidia's GPUs are excellent for training, but they have a critical weakness in inference: memory bandwidth. The bottleneck in inference is not compute power but data movement speed. GPUs must repeatedly fetch weights from VRAM, a process that is slow, expensive, and power-hungry.

Why Cerebras Excels at Inference

Cerebras's WSE-3 (third-generation wafer-scale engine) is a chip that covers an entire wafer:

ParameterWSE-3NVIDIA H100
Chip area46,225 mm²814 mm²
AI cores900,00016,896 CUDA cores
On-chip SRAM44GB
Inference speed15-20x H100Baseline

The WSE-3 is 57 times larger than the H100. This is not a marketing number but a key factor in inference performance: 44GB of on-chip SRAM places memory right next to compute cores, reducing data movement distance from centimeters to micrometers. That's where the inference speed gap comes from.

What's the trade-off? Low wafer-scale packaging yield, high cost, and a missing CUDA ecosystem — it's nearly unusable for training, leaving Nvidia's moat in training intact. But for inference? Cerebras's architecture has a structural advantage here.

OpenAI's Two-Pronged Strategy

OpenAI's relationship with Cerebras is not new. In January, OpenAI signed its first major contract: 750 megawatts of compute capacity over three years, valued at over $10 billion.

On April 17, The Information reported the follow-up: OpenAI added commitments bringing total spending over three years to more than $20 billion, potentially reaching $30 billion, with:

  • Equity warrants that increase with consumption, up to 10% ownership of Cerebras
  • A separate $1 billion investment to help Cerebras build data centers

This is not just a procurement contract; it looks more like a strategic alignment. At the same time, OpenAI is also developing its own custom ASIC chips in partnership with Broadcom, targeting mass production by the end of 2026.

OpenAI is consciously diversifying its compute supply — building a parallel track outside Nvidia while betting on both dedicated inference chips and its own custom silicon.

Nvidia's Countermove: Buying the Competition

Nvidia acquired Groq for $20 billion late last year. Groq's LPU (Language Processing Unit) is a pure inference architecture that, like Cerebras, bypasses the GPU memory bandwidth bottleneck. It was Cerebras's main competitor in the standalone inference chip space.

After the acquisition, Groq went from an independent player to a division within Nvidia. This is a defensive move, but its effect is direct: bringing a competitor in-house is cheaper and more thorough than fighting them in the open market.

The result is that Cerebras's main architectural rival in this space has disappeared, becoming Nvidia's own product. This landscape is favorable for Cerebras's customer expansion in the coming years.

Cerebras Renews Its IPO Push

On April 17, Cerebras refiled its IPO application with the SEC, targeting a valuation of approximately $35 billion, planning to raise $3 billion, and aiming to complete the listing in the second quarter.

The first IPO attempt failed because its largest customer, G42 (an Abu Dhabi AI company), was scrutinized by CFIUS. Cerebras converted all of G42's voting shares to non-voting shares, and CFIUS finally cleared the deal on March 31, 2025.

The conditions are different now:

  • Groq was acquired by Nvidia, removing its main competitor
  • OpenAI is an anchor customer, with a $20 billion commitment as the biggest asset in its roadshow
  • The valuation target jumped from $23 billion in the Series H round in February to $35 billion, a 50% increase before the listing

OpenAI, Nvidia, and Cerebras each have their own calculations in the inference battlefield, but they are all betting on the same thing: demand for inference compute will multiply several times over in the next five years. The cost of securing a position now is far cheaper than entering the market two years later.

Sources: OpenAI to Spend More Than $20 Billion on Cerebras Chips, Receive Equity Stake (The Information); AI chipmaker Cerebras files to go public after scrapping IPO plans last year (CNBC); CocoLoop, Two $20 billion deals: OpenAI and Nvidia are waging a "war of inference" (PANews); Cerebras IPO 2026: The $25B Nvidia Challenger (Nerd Level Tech)