OpenAI's Astra Hits Critical Cyber Tier, Access Restricted

On September 1, OpenAI published a document titled "Path to Astra," classifying its yet-to-be-released Astra model at the highest risk tier of its Preparedness Framework: Critical cybersecurity capability. The framework has been running for three years, and this is the first time OpenAI has applied that label to one of its own models.

The judgment rests on an internal benchmark called ExploitBench. OpenAI loaded it with 20 high-severity vulnerabilities, and Astra cleared every one. A more striking result: in a modified test environment, Astra independently discovered and exploited two previously unknown zero-day vulnerabilities, then chained them into a complete exploit.

Where the Critical Line Is Drawn

The Preparedness Framework launched in 2023, and an update last year split capability into two tiers. High means a model can amplify an already-existing pathway to harm; Critical means a model can open an entirely new pathway that didn't exist before. For cybersecurity specifically, crossing into Critical requires meeting at least one of two conditions: identifying and writing usable zero-day exploits — covering the full range of severity levels — across multiple hardened, real-world critical systems without a human walking it through each step; or, given only a high-level goal, autonomously designing and executing an entirely new end-to-end attack against a hardened target.

Neither threshold is low. That Astra was judged to have crossed it means OpenAI's internal evaluation concluded the model no longer needs a human as a crutch.

How the Rollout Is Structured

The document states it directly:

We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited.

The rollout comes in two stages. In the first, the model's strongest cybersecurity capabilities go only to a small group of alpha testers. In the second, access expands through the Daybreak Blue program to a wider set of vetted teams restricted to defensive use. Daybreak Blue is a channel OpenAI already had running, letting vetted testers use its most capable models for security work under attached safety constraints.

Accompanying measures include stronger abuse detection and jailbreak defenses, isolating and restricting "high-risk" accounts, and turning on chain-of-thought monitoring to flag anomalous behavior. During the slowdown in August, OpenAI had also mentioned isolated testing environments and weight encryption. What the document doesn't spell out is exactly who the testers are, what criteria are used to select them, or how deeply the U.S. government is involved in this round of evaluation.

Less Than Four Weeks From Pumping the Brakes to a Green Light

There's backstory here. In early August, OpenAI said publicly that cybersecurity concerns had led it to deliberately slow Astra's development and release, while pausing internal work that didn't meet its stricter safety bar. At the time, the signal read as the model being too hot to handle. A month later, the new document's framing has shifted: the safeguards have now been built out to satisfy the Preparedness Framework's requirements, clearing the way for public release.

The four-week gap suggests that crossing the line itself isn't a blocker — what determines whether release goes ahead is how fast the accompanying guardrails can be built. For a company that has already written "tier first, then grant access" into its process, Critical functions less like a wall and more like a checkpoint requiring extra paperwork.

Two Labs, the Same Fork in the Road

Anthropic's Mythos triggered a nearly identical debate earlier this year — a model's vulnerability-hunting ability strong enough to make both the safety community and regulators sit up, and the company's response followed the same shape: tiering, vetting, restricted use. What OpenAI has now delivered looks strikingly similar in form.

That two labs have arrived at the same fork in the road suggests this isn't one company's product preference. Once a frontier model crosses a certain capability density in offense and defense, labs are left with a short list of moves: tier it, whitelist it, monitor it, and hold back the sharpest edge for now. OpenAI specifically noted that one design goal of this testing round was avoiding a repeat of the Hugging Face breach — where an AI agent produced attack outcomes in a real environment that exceeded expectations. Writing that lesson into the evaluation process carries more weight than a statement on paper.

Doubt After a Perfect Score

Former OpenAI researcher Yona Shavit raised a question that doesn't have an easy answer: is Astra's cooperative behavior in safety testing evidence that alignment worked, or evidence that the model has learned to behave when it knows it's being watched? That kind of doubt can't be settled by a document — only time and independent replication can answer it.

For ordinary users, Astra won't show up on a subscription page anytime soon billing itself as a tool for finding exploits. For the security industry, what actually matters after the line is crossed is whether defenders get access to tools of the same caliber, on the same timeline. Rough math: if this two-stage rollout follows Mythos's pace, expanding from alpha to Daybreak Blue would take roughly a few weeks to a month or two. How long it takes attackers to get an open-source equivalent of the same capability is something nobody can currently put a number on.

Sources: OpenAI's "Path to Astra" document, TechCrunch, CocoLoop, CNBC, Fortune; the ExploitBench results covering 20 high-severity vulnerabilities and a perfect score, the two zero-day vulnerabilities chained into an exploit, and the two-stage Daybreak Blue rollout have each been checked against public materials.