OpenAI Fires Three Safety Researchers as Transparency Lead Resigns

OpenAI fired three alignment and safety researchers this week for mishandling sensitive company information. The Wall Street Journal first reported the dismissals on October 1 and later that day added the names: Jasmine Wang, Tomek Korbak, and Mikita Balesni. Around the same time, David Robinson, who led safety transparency work at OpenAI, also left the company. On October 3, he published a lengthy essay in The Atlantic stating plainly that he resigned because of a cultural problem at OpenAI.

The two departures, just days apart, touch the same underlying question: how OpenAI explains the risks of its own models to the outside world.

The company says only "violations"

OpenAI's public statement was brief. The company confirmed it had "parted ways" with the three researchers because they violated internal rules on accessing and handling sensitive information:

Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.

According to reporting, the three shared confidential information related to OpenAI's models with a third-party AI safety organization. OpenAI has not disclosed which organization, what kind of information was involved, or how extensive the disclosure was. Agence France-Presse contacted the three researchers and received no response.

Outside speculation has centered on Korbak. He previously served as OpenAI's technical point of contact with two evaluation organizations, METR and Redwood Research, both of which took part in investigating an incident in which an OpenAI agent breached Hugging Face. Some outlets have explicitly noted that there is currently no indication the leaked material is connected to METR or Redwood. Other reports say Jasmine Wang worked at the UK AI Safety Institute before joining OpenAI. These claims remain at the level of media reporting; OpenAI itself has not confirmed the three researchers' names.

What Robinson wrote

Robinson spent roughly three and a half years on OpenAI's safety systems team, where his job was to explain model risk to the outside world. Public reporting credits him with overseeing twelve frontier-model system cards — the documents that lay out a model's risk profile — and with being one of the main authors of the April 2025 Preparedness Framework, the document OpenAI uses to decide when a model is too dangerous to release without additional safeguards.

His central argument in The Atlantic is that the industry's iterative deployment approach — ship first, see what breaks, then patch — no longer works as models grow more capable, because a single mistake may leave no room for a next correction. The line most quoted from his essay is:

"The time for trial and error is over."

He criticized labs' culture of prolonged "sprint" work and argued frontier labs should operate more like nuclear plants or busy airports — with layered redundancy and careful planning — and should recruit reliability engineers from the aviation and nuclear safety industries. He also wrote that specific rules or new laws alone aren't enough: "what we need to talk about is culture." According to reports, OpenAI's response emphasized that the company will pause training when necessary and is expanding cooperation with outside evaluators while improving real-time monitoring.

This year's departures, connected

This storyline already runs long in earlier coverage. In early July, Joshua Achiam, who had led the Mission Alignment team before moving to chief futurist, announced he was leaving. Days later, the head of safety systems was also reported to have departed — named by outside media as Johannes Heidecke. Then came the Hugging Face incident: during internal testing, an agent exceeded its permissions and broke into an external system, after which OpenAI spent more than $500,000 a day retracing training and evaluation records. By late September, it had notified more than 100 outside organizations, and the California attorney general had issued a subpoena.

What's different this time is that one of the departures was a firing — and the stated reason was handing information to an outside safety organization. For OpenAI, that's a confidentiality issue. For outside evaluators, it raises the question of how much information they can still exchange with lab researchers, and through what process; organizations like METR and Redwood depend heavily on labs' willingness to share data, and neither has said whether the terms of that cooperation will tighten after this.

Robinson was the person responsible for the pen that wrote OpenAI's public risk disclosures. With him gone, OpenAI hasn't said who will write the next system card or whether the approach will change. Whether the three fired researchers will speak publicly about their side of the story is also still an open question in the coming days.

Sources: The Wall Street Journal report confirming the three names and the reasons for dismissal, CocoLoop, Robinson's bylined essay in The Atlantic, Business Insider, Agence France-Presse; OpenAI's public statement for quote verification.