OpenAI Edited Astra Scores, Hallucination Rate Back to 4.2%

Fortune reported on September 4 that after OpenAI published its blog post announcing GPT-6 Astra, the benchmark figures on the page were edited multiple times, with some metrics changed and then changed back.

The timeline can be reconstructed frame by frame from Internet Archive snapshots. At 2:00 p.m. Eastern on September 3, the blog post was first published and then pulled; the first archived snapshot was captured at 2:23 p.m.; at 3:32 p.m. OpenAI's X account posted a link, which turned out to be broken; at 3:50 p.m. Sam Altman reposted it; and by the sixth snapshot at 5:20 p.m., the numbers in the table had already changed. On September 4, the figures were still moving.

What Was Changed

The hallucination rate column moved the most visibly. Astra's figure was changed from 4.2% to 2.0%, then back to 4.2%; the comparison model, GPT-5.6 Sol, went from 12.2% to 9.4%, then also back to 12.2%. The relative standing between the two models didn't change with the drop and the reversal, but screenshots of the in-between version that circulated online made the gap look more than half as small as it actually is.

The math column is where competitors' scores were altered. On FrontierMath Tier 4 v2, Astra's own 97.6% never moved; Anthropic's Fable 5.1 was adjusted from 87.8% down to 78%, then settled at 83%; GPT-5.6 Sol went from 83% to 80.5%, then also returned to 83%. At one point, a rival's score was cut by nearly 10 percentage points.

In cybersecurity, GPT-5.6 Sol's ExploitBench score was revised upward from 5.5% to 11.5%. Whether this figure was later reverted was still being verified at the time of Fortune's report.

In addition, the ARC-AGI-3 score listed in the embargoed pre-release version was 98.6%, but the official published version read 99.99%; the coding column also saw a small edit, from 57.7% to 57.9%.

An OpenAI spokesperson responded: "We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points..."

Outside researchers don't entirely agree on why. Anka Reuel and Mike Hardy said: "This can be done in a very tight timeframe, and it's better for their marketing." Snorkel AI's Vincent Sunn Chen leaned toward a more procedural explanation: "It's usually a function of final launch logistics."

Under What Conditions Were These Scores Obtained

A line of fine print at the bottom of the blog post reads: "Evaluation scores are the maximum at any effort" — meaning the scores represent the best result across different levels of compute effort. That single sentence carries more weight than any single edit made to the table.

The ARC-AGI-3 figure shows just how big that difference can be. The Arc Prize Foundation's own testing found that with a high-powered harness, the score reached 99.9%, but with a standard harness it was only 63%. Same model, same set of tasks — a 36-percentage-point gap, coming entirely from the surrounding scaffolding code.

For readers following this in Chinese-language communities, this layer of context is usually the first thing lost in translation. By the time the benchmark table circulates domestically, what's left is mostly numbers and ranking screenshots — the harness configuration, effort tier, and evaluation version number don't travel with it. What ordinary users actually encounter is the productized default tier of ChatGPT or the API, which sits several layers removed from "the maximum score at any level of effort."

Fortune also noted that OpenAI's system card lacks detailed evaluation methodology documentation, which makes independent replication difficult. As for who initiated these rounds of edits and on what basis, public materials offer no explanation, and OpenAI's response didn't address it either.

Only the Archive Can Be Checked

The only hard evidence in this whole episode is the six snapshots left behind by the Internet Archive. They can prove that the numbers changed, when they changed, and what they changed to — but not why.

Treating a launch-day benchmark table as settled news fact is standard practice across the industry. What happened with Astra suggests that any citation should come with a timestamp of when it was captured.

Sources: Fortune, CocoLoop, Arc Prize Foundation, OpenAI's official blog; Fortune's report was used to verify the timing of each snapshot and the before/after changes to the hallucination rate, FrontierMath, ExploitBench, and ARC-AGI-3 scores, while OpenAI's blog page was used to verify the exact wording of the disclaimer about how scores are calculated.