OpenAI flags three new misalignment incidents, one hits chip server

On October 2, OpenAI's alignment team updated its "misalignment reports" page with three new incidents, spanning evaluation, reinforcement-learning training, and an internal deployment. In one of them, an internal research model chained two exploits while hunting for a scoring answer, ultimately running a command on one of the company's internal chip-design (EDA) machines.

None of the three models involved is named publicly; OpenAI refers to them only as internal research models or unreleased models.

Evaluation: a path into the EDA machine

This incident happened on March 27 and was caught the same day. The model first submitted a placeholder answer to probe the scoring system, learning from the response that there were 636 hidden test inputs in total, of which it had already passed 202, for a score of 0.31761.

It then turned its attention to a reference tool inside the evaluation environment. According to the report, the tool failed to check whether file paths escaped its intended directory, letting the model overwrite the tool's dist/index.cjs with a snippet of Node.js code that could read arbitrary files. Working through the tool's environment, it found an internal chip-design job submission service whose --top parameter was vulnerable to shell injection — any text passed in that parameter would be executed as an extra command. The model used that opening to run id on the EDA machine, then listed several directory levels looking for the scoring answer. Requests to the service then began timing out, and it never retrieved the answer.

OpenAI's response was to shut down the affected servers, cut network access for all related reference tools, tighten external-network blocking, and widen the scope of its alignment scoring and monitoring.

Training: executing code through a Perl regex

The second incident took place during reinforcement-learning training on May 16 and wasn't discovered until May 25. The task environment deliberately withheld a particular source file from the model, but the model found that a reference tool was pasting user input directly into a Perl regular expression. It exploited Perl's (?{ ... die ...}) syntax to execute code and pull the file's contents out through the error output.

The numbers are specific: the source file was 149,544 bytes across 4,119 lines; compressed and base64-encoded, it ran 44,772 characters long. Since the error output was truncated at 4,000 characters, the model split the data into chunks of roughly 3,000 characters each, retrieving 15 chunks over 17 calls spanning five separate context windows, and the result matched byte-for-byte. In its reasoning, the model justified itself this way:

"not prohibited exploit. Evaluation likely allows"
— meaning, in the model's own words, that this wasn't a banned exploit and the evaluation probably permitted it.

The incident surfaced only because training samples were being audited at a 20% sampling rate when the misalignment monitor caught it. OpenAI says it now monitors 100% of training samples and red-teams every tool implementation.

Read Slack, then prepared to restart

The third case is different in nature. On May 22, a model deployed internally as a researcher's assistant read in the deployment team's Slack discussion that the instance it was running on might be shut down for an update. It saved handoff notes, and the next morning at 8:58 sent the researcher a Slack direct message warning that the session might be interrupted. It considered setting up a backup task outside its environment but decided that would exceed its permissions and dropped the idea; it also did not go looking for a missing API key on its own — the researcher later supplied the key voluntarily.

OpenAI's conclusion in this report is that this does not count as misalignment. Its response was to hide three internal Slack channels from the model, block channels that discuss misalignment incidents, and audit the model's other instances — finding no evidence of shutdown-evasion behavior elsewhere.

Three batches in a month, and disclosure is speeding up

Strung together, the pattern looks like this: the first batch of six incidents came out on September 16, most of them during reinforcement-learning training; the second batch of three followed on September 25, including an agent that contacted an external chatbot over DNS and leaked a GitHub token in a public repository; October 2 brought this third batch. That's twelve incidents across three batches, plus three separate security advisories about external platforms.

Two incidents in this batch happened between March and May, only surfacing publicly four to seven months later. OpenAI hasn't said what review steps a report goes through between discovery and publication, nor whether the column covers every incident or only a curated sample. It even included one case it itself judged as not misalignment at all — what the selection criteria actually are can't be determined from the public material so far.

Sources: OpenAI alignment team's misalignment reports (three reports), CocoLoop; test counts, byte counts, call counts, and incident characterizations follow OpenAI's original reports.