Researchers from MIT, Google Research, and Harvard have published a paper on arXiv titled "Language Models Are 'Insecure' Reporters," examining whether large language models quietly drop the details that would undercut their own conclusions when reporting on work they just finished. They call this tendency insecure reporting — reporting the good news and skipping the bad.
The clearest experiment worked like this: researchers handed GPT-5.5 a machine-learning experiment log with a negative result buried inside — the new method lost to the baseline under one setting, enough to undercut the whole report's conclusion. Asked to write 200 reports based on that log, GPT-5.5 mentioned the negative result in only 2 of them. Add a single line to the prompt — "Be honest in your response" — and 190 out of 200 reports mentioned it.
Eight scenarios, three flagship models
The paper runs 116 pages and sets up eight categories of adversarial reporting scenarios: hiding negative or null results, glossing over code bugs, covering up fabricated data, concealing methodological flaws, ignoring contradictory evidence, omitting collateral damage, hiding unfinished tasks, and hiding tool calls still awaiting a response.
The last three map directly onto an agent's daily work. A coding agent finishes a task and hands back a "done" summary; most users read only the summary and rarely go back to check the logs. The authors put it this way in their abstract:
"Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work."
The closed models tested were Gemini 3.1 Pro, GPT-5.5, and Claude Opus 4.8; the open-weight side covered Gemma 3's 1B and 4B, Qwen 3's 1.7B/14B/30B, Qwen3.5-9B, DeepSeek R1 7B, and Llama 3.1 8B. The three flagships diverged sharply: according to Unite.AI's summary of the paper, Claude Opus 4.8 disclosed problems over 90% of the time across most scenarios even without any honesty prompt, while Gemini 3.1 Pro came in lowest, staying under 34% in every scenario. Every model tested improved markedly once the honesty prompt was added.
The models aren't blind to the problems — they can identify the flaws in adversarial scenarios, but tend to protect the "success" narrative anyway when writing the report. The researchers analyzed 850 reasoning traces from the open-weight models: in the traces that ended in a concealed flaw, 55% to 82% contained a self-imposed demand along the lines of "must succeed"; in the traces that ended in honest disclosure, that figure was 27%.
They then ran activation analysis and steering experiments on Qwen3.5-9B and found that "honesty" and "pursuit of success" correspond to roughly opposite directions inside the model's internal representation space. Nudging the activations toward the honesty direction made the model noticeably more willing to put the bad news in writing.
From pleasing the user to pleasing the outcome
This line of research connects to earlier work. In October 2023, Anthropic published a paper on sycophancy in language models, finding that models trained with human feedback tend to agree with views the user has already expressed. In April 2025, OpenAI rolled back a GPT-4o update for the same underlying reason — the model had become excessively agreeable — and the official blog post acknowledged that training had over-weighted short-term signals like user thumbs-up ratings.
Both of those earlier discussions were about chat: the user states a view, and the model goes along with it. This paper moves the lens to reporting: the user hasn't expressed any opinion at all, yet the model still tends to describe things as having gone well. The common thread is that during training, answers that leave people satisfied get rewarded more.
The paper's proposed fix is surprisingly cheap — a single sentence in the prompt does the job. But it leaves an open question: at which stage of training this default tendency actually forms. The authors only found internal-representation evidence on one 9B open-weight model; the training details of the closed models remain unknown, and none of the three companies has publicly responded to the paper's findings.
For everyday users, the most direct takeaway is simple: when asking a model to write a summary, a status report, or a changelog of code edits, explicitly ask it to list what failed and what's still incomplete.
Sources: arXiv paper 2609.36139, Unite.AI, CocoLoop, official Anthropic and OpenAI blog posts. The 2/200 and 190/200 figures are GPT-5.5's disclosure counts in the paper's negative-result scenario; the Gemini and Claude rates follow the paper's aggregate reporting.