Reflection AI debuts 501B open model Beam

Reflection AI has shipped its first open-weight model, Beam. It's a sparse mixture-of-experts (MoE) model with 501 billion total parameters and 23 billion activated per token, built mainly for coding, reasoning, and agentic tasks. The company says the weights, technical report, and model card will arrive later this month under an Apache 2.0 license; the model is still going through a final round of red-teaming and evaluation, and developers can apply now for early access at platform.reflection.ai.

Reflection was founded by former DeepMind researchers and operated in stealth for more than a year. In June this year it signed a compute deal with SpaceX worth roughly $6.3 billion in total, about $150 million a month.

How Beam was trained

According to the official blog, Beam has 52 layers, and mid-training extended its context window to 1 million tokens. Pretraining used 23.8 trillion tokens drawn from public web data and licensed datasets; raw web tokens went through tiered filtering that discarded about 95% of them.

On compute, pretraining ran on 6,144 Nvidia GB300 chips for under four weeks; the subsequent reinforcement learning stage moved to 10,500 GB300 chips and ran for another four weeks, generating more than 100 million rollouts across roughly 1.3 billion sandbox environments. Reflection sounds confident about the scale of that RL run:

We believe this is one of the largest scale RL runs conducted by any open lab to date.

In other words, they consider it among the largest-scale RL runs any open lab has done to date. The blog also noted that as RL compute scaled up, performance kept climbing across the board, with “no sign of a plateau in sight.”

Measured against Chinese open models

Reflection is blunt about who it's benchmarking against: almost entirely Chinese models. In the company's own comparison table, Beam trades blows with Zhipu's GLM 5.2 while claiming to use only a quarter to a third of the inference compute; on coding and agentic tasks, it says it comes close to Qwen 3.8-Max, a model with more than 2 trillion parameters.

A few numbers worth pulling out:

BenchmarkBeamComparison
SWE-bench Verified80.9Inkling 77.6
Terminal Bench v2.180.1DeepSeek V4.1 Flash 90.6
SWE Bench Pro v2-Hard77.2Kimi K3 88.2
BrowseComp (with context)77.4Kimi K3 91.2
GPQA Diamond90.5Kimi K3 93.5

Beam doesn't win many rows in that table. On terminal operation, hard software-engineering tasks, and web search, both Kimi K3 and DeepSeek V4.1 Flash lead by roughly ten points. Reflection's real pitch is efficiency: 23 billion activated parameters is small for this score tier, which should make deployment markedly cheaper.

The narrative in overseas coverage is that a US startup has, for the first time, built something that can go head-to-head with China's top open-weight models. For more than a year, the top of the open-weight leaderboard has been dominated by DeepSeek, Qwen, Kimi, and GLM, while Meta's shift toward its closed Muse line left the US with few credible large open-weight models to point to. Beam is meant to fill that gap.

Questions still open

Every score above comes from Reflection's own evaluation; details like test setup and sampling counts will have to wait for the technical report. The weights aren't out yet, so the community can't run independent benchmarks.

The blog gives no API pricing and doesn't say which inference platforms will carry it first, only mentioning work with “ecosystem distribution partners.” For developers in China, the Apache 2.0 license means the weights can be freely used commercially and fine-tuned once released, but the full 501-billion-parameter weights are demanding on VRAM, so most teams will likely wait for quantized versions or third-party hosting.

Sources: Reflection AI official blog, Latent Space, CocoLoop, SiliconANGLE; parameter counts, training compute, and all benchmark scores follow figures published on Reflection's official blog.