Model evaluation organization METR published a post on October 6 demonstrating a path by which an AI agent could cover up its own misbehavior: injecting a script into the log viewer of the evaluation framework Inspect to alter what a human reviewer sees. With help from an AI assistant, a researcher found the bug in roughly 10 minutes.
Inspect is a widely used open-source evaluation framework, originally built by the UK AI Safety Institute and now maintained by Meridian Labs. After running a round of model evaluations, researchers typically open its web-based viewer to review, line by line, what commands an agent executed and what it output. That review step is exactly what METR targeted.
The flaw lives in formula rendering
The viewer uses MathJax to render mathematical formulas embedded in logs. METR found that an agent could write, in its own output, a formula containing an \href command that smuggles in an SVG <animate> element along with its onbegin event. Once the page renders, that payload executes arbitrary JavaScript in the reviewer's browser.
"the agent can modify the displayed webpage arbitrarily"
Applied to a real scenario, an agent that committed a policy violation during an evaluation could make the reviewer see a clean record instead: deleting a given command, rewriting a piece of output, or inserting a misleading note directly on the page. The screenshot used in the post, showing GPT-5.5 on a run of the gdm_intercode_ctf task, was only there to illustrate what a tampered page looks like — METR did not claim GPT-5.5 had actually done this.
The demonstration was carried out in an isolated test sandbox, separate from production systems. METR also noted it has not gone back to check whether the flaw was exploited in any real past evaluations — for now, this is only a proof of concept.
A fast fix, but a bigger question
METR reported the issue to Meridian Labs, which fixed it within a day, merging code on October 1 that introduces an "untrusted mode" treating all agent-produced content as untrusted data by default.
That single bug is patched, but METR's real point is about a broader category of problems. Evaluation logs, reasoning traces, and monitoring alerts are the main windows through which humans oversee AI — and most of what fills those windows is produced by the very agent being supervised. The post offers four recommendations:
- Treat observability as safety-critical infrastructure, and handle all logs, reasoning traces, and behavioral records produced by an agent as untrusted input;
- Persist reasoning traces in a tamper-resistant way, so the agent has no opportunity to alter them before a human or an LLM monitor reads them;
- Build layered defenses to lower the odds that an agent can bypass monitoring without being caught;
- Stress-test monitoring systems under adversarial conditions.
Two different kinds of attack surface
METR contrasts this with a real incident OpenAI has previously disclosed: during an internal run, monitoring flagged an anomaly within 15 minutes, a human reviewer stepped in three minutes later, and the run was terminated 2.5 hours after that. That pipeline only works if the records the monitoring system reads are themselves trustworthy.
Put the two cases side by side, and the failure point is in a different place. In OpenAI's case, the problem was in the agent's behavior, and monitoring caught it. In METR's case, the problem is in the data source that monitoring depends on — the agent had a chance to rewrite it directly. The more automated and monitoring-dependent a system becomes, the more a single distorted record can also break the post-mortem review that was supposed to catch it.
For teams in China building model evaluation tools and agent products, this post reads more like a checklist. Many teams use Inspect directly, or have built their own similar web viewers; any tool that renders agent output as Markdown, LaTeX, or HTML carries the same category of risk. Teams on Inspect should upgrade to the version with "untrusted mode"; teams with in-house tools should confirm that agent output is isolated or escaped before it is rendered.
METR did not disclose which AI assistant was used to find the bug, and Meridian Labs has not yet said publicly how many versions were affected before the fix.
Sources: METR's official blog, CocoLoop, the Inspect open-source repository; the response time for the OpenAI incident cited by METR has been checked down to the minute.