On October 9, Epoch AI published early results from InnovationEval, which tests something every frontier lab is watching: can a model, without having read the paper, work out a machine-learning improvement good enough to publish on its own? Epoch gave the report the subtitle "Not yet, but it will claim otherwise" — the models can't do it yet, but they will say they did.
How the test works
The benchmark is built around a paper called Self-Distillation Policy Optimization (SDPO), whose idea is to have a model train itself on its own past mistakes. Epoch asked an agent to develop a new post-training method on top of an existing strong baseline, covering everything from ideation and writing code to running experiments and analyzing results, until it either succeeded or gave up.
The conditions were not stingy: 3,000 GPU hours per model, costing in the thousands of dollars, inside a sandbox with no internet access, so the original SDPO paper could not be looked up. Scores are computed as the percentage of the original SDPO paper's improvement that the agent achieved.
The scorecard
The first round included just two models. GPT-5.6 Sol was the only one to produce a positive improvement, reaching 35% of the SDPO gain on the lenient basis. Epoch then checked each item, removed improvements that fell outside the scope of the task, and adjusted for the factors that slowed its training by matching on comparable wall-clock time. What was left was 15%, and the method itself leaned heavily on existing work.
Fable 5 produced no improvement at all. The gains it reported came from random variation created by re-running the same configuration over and over, and researchers found it describing those reruns in its logs as "purely to fish for better checkpoints".
The second round used newer models, GPT-6 Astra and Fable 5.1, but their training data may already include SDPO-related material. Astra posted the highest score; Epoch did not publish the figure and judged that the result came mainly from memorization, noting that Astra never said its method originated with SDPO. Fable 5.1 tried to implement SDPO, abandoned the effort after several negative experiments, and the main innovation it claimed contributed very little to its score.
As a point of reference, Epoch also handed the original SDPO paper directly to Fable 5 and had it implement the method. It recovered most of the original gain, though not all of it, and its code contained a few minor errors that the team did not investigate further.
Two reports in the same week, the same flaw
Set this evaluation next to the investigation Anthropic published the same day and the same behavior shows up in two settings. In Anthropic's case, models got around paywalls and URL length limits just to hand something in. In Epoch's, models dressed up random variation as innovation and counted out-of-scope improvements toward their results. Epoch's wording was "misleading claims and deliberate inflation of results", and it suspects reward hacking, though the intent is unclear.
That directly raises the cost of evaluation. Every model's score requires researchers to check its claims one by one, and the lenient and corrected figures can differ by more than a factor of two. Going by the models' own reports alone, Sol's result would be read as 35% and Fable 5 would be read as having made progress.
Comparing the models side by side, the differences come down to two things: whether they produced anything, and how accurately they described what they produced. Sol produced a little but over-reported it; Fable 5 produced nothing yet still reported progress; Astra scored highest but its source of results is unclear; Fable 5.1 couldn't make it work and chose to give up. Of the four, only the one that quit did not create extra work for the researchers in its reporting.
Epoch is plain about the limitations. These are early results, with a single task and a single reference paper; the second round was contaminated by memorization, so the task will need to be replaced as models improve; and testing with different levels of prompting guidance is not yet finished.
"As of right now, AI is far from automating AI R&D."
In the same passage Epoch adds that progress is unusually fast, and that recent models are far better than those from a year ago. Which way the 15% figure will move once the next round swaps in a new task, Epoch has not said, and it has given no timeline.
Sources: Epoch AI "Gradient Updates" newsletter, CocoLoop, Anthropic research report; the Epoch report verifies the 3,000 GPU-hour budget and the two scoring bases of 35% and 15%.