In a report published on September 6 titled "Research acceleration: The view inside OpenAI," the company said it has hit its internal "automated research intern" milestone. The report defines that milestone as a system that can complete well-specified research tasks under human direction, including tasks that would take a skilled researcher days to finish.
The claim rests on a set of internal usage data. As of mid-August, AI agents were logging the equivalent of 3.1 agent-workdays for every standard eight-hour workday put in by a human researcher. Over the same period, the median researcher's daily inference spend topped $600, and the 90th-percentile user topped $7,000 a day. The number of researchers running four or more agents at once has been climbing. The report also notes that experiments per active researcher in August 2026 hit the highest level since OpenAI began tracking the metric in January 2025.
The limits the report spells out itself
The document doesn't oversell the result. Success rates are only tallied for tasks with a clear ground truth; research work without an objective right-or-wrong answer isn't included. For tasks in the four-to-eight-hour complexity range, more than half still require a human to step in to finish the job. The report states directly that "the measurements are preliminary," and warns that various potential bottlenecks could keep overall research throughput from matching the gains shown in any single metric.
On the endgame, the report is blunt: "We do not yet know how to safely get all the way to aligned, full RSI" — the company is acknowledging it hasn't worked out a safe path to full recursive self-improvement.
The 3.1 figure can't currently be verified from the outside. It depends on OpenAI's internal method for converting activity into "agent-workdays," and that conversion methodology hasn't been made public.
The last timeline came eleven months ago
Neither milestone date is new. In an October 2025 livestream, Sam Altman said:
"we think it is plausible that by September of next year, we have an intern-level AI research assistant and that by March 2028, we have a legitimate AI researcher."
By OpenAI's own account, the first milestone landed in the month it was predicted for. The second is still pegged to March 2028; the report describes it as an automated AI researcher that can advance deep learning and alignment research under human oversight and support iterative improvement, rather than a system limited to narrow tasks.
Two less tidy data points from this same timeline made it into the report as well. On July 20, the company's container service was temporarily shut down following a security issue. On August 7, Astra was flagged as having reached critical cybersecurity capability. The same internal tooling that's driving up research throughput is also generating security issues that need to be handled.
The bill is climbing too
The inference-spend numbers are easier to sanity-check than the 3.1 figure. A $600-a-day median per researcher means a hundred-person research team would burn through roughly $60 million a year on agent calls alone, based on a rough 250 working days a year and excluding training. For the 90th-percentile users at $7,000 a day, a single researcher could run up close to $2 million a year.
That money behaves differently from training compute. Training is a one-time outlay that produces a model you can reuse indefinitely; researchers running agents for experiments is pure consumption — spent and gone — and the more fluently people use it, the more they spend. The report's observation that more researchers are running four-plus agents at once points directly at the slope of that spending curve.
For outside observers, March 2028 is the next date to check against. Until then, the only public material available for cross-checking is the report OpenAI itself has published.
Sources: OpenAI's official research acceleration report, CocoLoop, unite.ai; verified against the OpenAI report page for the 3.1 agent-workday ratio, median and 90th-percentile inference spend, the share of four-to-eight-hour tasks needing human intervention, and Altman's original wording on both timeline milestones.