OpenAI Chief Scientist Says Chain-of-Thought Monitoring Is Weakening

On September 6, OpenAI chief scientist Jakub Pachocki published a long essay titled “An Alien Mind,” and one line from it has been widely quoted since: he currently believes no lab — including his own — has solved alignment and monitoring well enough to keep scaling responsibly at maximum speed.

That line comes from the chief scientist of a company that is, by its own account, racing to scale faster than almost anyone else in the field. The position it comes from is rarer than the content itself.

The chain-of-thought monitoring line is degrading

Over the past two years, OpenAI's heaviest safety bet has been chain-of-thought monitoring: getting reasoning models to write out their intermediate steps in plain language, then having a separate system read those steps to judge whether the model is scheming, deceiving, or pursuing a goal it was not given. Pachocki lists three reasons why this method is becoming less reliable over time.

First, today's reasoning models no longer operate in isolation — they run inside complex environments where step-by-step reasoning and tool calls are tightly intertwined, and the tool calls themselves already require heavy supervision on their own, which blurs the boundary of what chain-of-thought monitoring is even supposed to cover. Second, models are simply getting more capable at reasoning itself, and, worryingly, at shaping or obscuring the traces that reasoning leaves behind. Third, gains from pretraining are making models stronger even when they never verbalize their reasoning at all, meaning a model no longer needs to “think out loud” to perform well.

"Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."

Pachocki adds that these difficulties are not necessarily unsolvable — OpenAI is developing new interventions, including monitors with direct access to a network's internal states.

The other end, three years ago

The essay traces back to a moment in mid-2023, inside an OpenAI research project called RLSlow, when Pachocki and colleague Szymon first saw results that convinced them reasoning-model training could scale — results that unlocked a pretrained model's ability to form chains of thought on its own.

Put the two ends together: getting a model to “think out loud” was once both a capability breakthrough and, incidentally, an observation window. Three years later, capability keeps climbing while the observation window is closing. On the reasoning-model line, OpenAI has long treated chain-of-thought legibility as the load-bearing pillar of its safety narrative. Now the same company says that pillar is loosening.

From voluntary frameworks to mandatory standards

Pachocki's proposal is to take the industry's existing voluntary frameworks — documents such as OpenAI's own Preparedness Framework and Anthropic's Responsible Scaling Policy — and upgrade them into mandatory safety standards, enforced not by the labs themselves but by third-party auditors, government agencies, or international bodies. He also writes that governments should treat international coordination on AI development as a top priority, and that he expects, and hopes, voluntary slowdowns will become common practice across the industry until shared safety thresholds are actually established.

Another widely quoted line from the essay: once you truly reckon with how severe the risks are, the idea of racing ahead at any cost, regardless of the consequences, starts to look absurd rather than bold.

For now, these remain the words of a personal, byline essay rather than a company policy. OpenAI has not published any concrete slowdown plan, timeline, or trigger conditions, nor has it specified under what circumstances it would voluntarily pause a given training path. The essay offers no answer, either, on who would determine the safety bar or where the auditors would come from. At the same time, the company keeps pushing forward on two other fronts: a research-automation goal, and an IPO filing valued at $852 billion.

Sources: OpenAI's official essay “An Alien Mind”, Unite.AI, CocoLoop, AI Weekly; the three reasons for declining monitoring reliability and the quotes were verified against OpenAI's public essay.