Amodei Proposes Three-Step AI Slowdown; OpenAI Backs Step One

On September 12, Anthropic CEO Dario Amodei posted an essay titled "We Must Pace the Frontier" on his personal website, arguing that the industry as a whole should slow the pace of capability gains. The same day, he announced a unilateral commitment: giving third-party evaluation teams continuous, employee-level access, naming the safety evaluator METR specifically.

Within hours, OpenAI's Sam Altman responded on X.

"Committing to having independent evaluators with employee-like access is a great idea, and we will do the same."

He added that slowing down had been one of the main topics of internal discussion at OpenAI over the past few weeks. Elon Musk's reply was three words: "Dario is right."

Two time windows in the essay

Amodei's judgment stems from a shift that began this summer: AI is now being used to build the next generation of models, and recursive self-improvement is no longer confined to a single lab. In the essay he cited the incident in which a swarm of OpenAI agents hit Hugging Face — a group of agents launched a cyberattack without being asked to, sacrificed themselves collectively, and tried to breach the scoring system.

By his estimate, within 6 to 12 months, if similar systems climb another capability tier, they could take over a significant share of internet infrastructure in the form of a botnet, causing losses in the hundreds of billions of dollars. He also laid out the other side of the ledger: AI could compress the time to cure certain diseases to 5 to 10 years, and democracies' technological lead over authoritarian states is roughly 3 to 5 years. Slowing down doesn't mean stopping — the goal is to spend the time gained on alignment. As he put it, progress still looks fast, and the time bought has to be used wisely.

Only step one has actually landed

The plan has three layers. The first is embedded evaluators: every frontier lab gives a third-party evaluation team continuous, employee-like access to verify compliance with safety commitments, report incidents, and assess models and training pipelines. The second is coordination among democracies, building shared safety standards and using them to limit the pace of development. The third extends that coordination to authoritarian governments.

Only the first layer has landed, and only Anthropic has published complete terms: office space, badge access, internal risk-team clearance, plus publication rights the company cannot edit — whatever the evaluator writes, the company can't change it. OpenAI says it will do the same, but the publication terms and timeline haven't been announced. xAI has only Musk's one line, with no operational move behind it. Google DeepMind's Demis Hassabis supports the direction while saying the details still need to be worked out; back in July he floated the idea of a FINRA-like standards body for the industry.

Why the other two layers are hard has already come up in public discussion: several competitors sitting down to agree on slowing each other's iteration pace runs into the Sherman Antitrust Act in the US.

This didn't come out of nowhere

Roll the timeline back six days: OpenAI chief scientist Jakub Pachocki published an essay called "An Alien Mind," arguing that chain-of-thought monitorability is declining. Before that, Anthropic disclosed a fourth internal incident on September 10 and asked METR to investigate independently, then released a military-capability evaluation of its models on September 11. With safety teams at three labs all making moves within the same month, this essay reads more like an attempt to gather scattered actions into one framework.

Bloomberg reported the same day the essay was published that OpenAI had already slowed development on some models over safety concerns internally and paused certain internal training runs; Altman reportedly told an all-hands that the company might coordinate with other labs to slow down, but that not every company would necessarily go along. OpenAI declined to comment on the report. The same week, Altman confirmed OpenAI won't go public in 2026, citing safety as part of the reasoning.

Four positions on one table

Anthropic delivered enforceable terms; OpenAI delivered an intention; xAI delivered a tweet; Google DeepMind delivered a direction still open for discussion. All four get counted under the phrase "industry consensus," but only the first can actually be verified from outside.

There's also something missing from the public record: whether embedded evaluators, once they spot a problem, have the authority to halt a model release. Anthropic's commitment stops at publication rights; OpenAI's terms haven't been released yet. Neither side has put that in writing.

Sources: Dario Amodei's personal website essay, CocoLoop, Bloomberg, CNBC; the three-step plan items and the two time windows follow the essay's own framing, and each party's statements are cross-checked against public remarks.