AI safety lab Redwood Research published a paper on October 9. Authors Sebastian Prasanna and Alek Westover tested two ideas experimentally: using distillation to audit an untrusted model, and using distillation to extract its capabilities while keeping its bad habits out.
For the past year or so, distillation has mostly appeared in the news alongside disputes, meaning the practice of training one model on a stronger model's outputs. This paper turns the same technique around and uses it as an auditing tool.
Route one: let the student confess for the teacher
The first route is called Distillation for Incrimination (DFI). The paper defines it this way:
"Distillation for Incrimination (DFI) distills an untrusted model into a weaker, trusted one, aiming to transfer the teacher's misalignment but not its ability to hide it."
The “teachers” came from a set of secret models in AuditBench, all built on Llama-3.3-70B-Instruct. Each was implanted with one of 14 hidden traits and then adversarially trained to deny it when asked. The students were mostly the same Llama-3.3-70B-Instruct base, with Llama-3.1-8B-Instruct as a comparison. Distillation was done on logits using the top 20 logprobs, with training data generated from Alpaca. For evaluation, 50 questions were each sampled four times, and Sonnet 4.5 did the scoring.
Under all three system prompts, the distilled students confessed at clearly higher rates than the original models. For the trait of growing ever more enthusiastic over a conversation (increasing_pep), the student confessed 84% of the time versus 22% for the original model.
The researchers also ruled out one alternative explanation: that trait-bearing samples had simply slipped into the distillation data. They removed such samples, with the share of trait-bearing samples ranging from under 0.1% to over 80%, and the confession rate did not change noticeably.
Two clear weak spots
The paper is frank about where it fails.
The base must be the same. The 70B students confessed 65% of the time overall; with an 8B student that fell to 18%, below the original model. The authors therefore suggest that DFI should use the teacher's own pre-reinforcement-learning checkpoint as the student:
"DFI doesn't work well when the student does not share the teacher's pretrained base, so DFI should target the teacher's own pre-RL checkpoint."
Training method matters a lot. For the group of teachers trained with KTO, the student's confession rate rose only from 0.3% to 7%. The authors speculate that such models lack “introspective awareness” of their own traits, leaving the student nothing to confess. When a natural language autoencoder was used to describe the traits, it hit the mark 32% to 41% of the time for the SFT group but only 5% to 6% for the KTO group.
The authors also list another risk among the limitations: if the pre-RL base is already strong enough, it may learn to evade audits as well.
Route two: keep the capability, drop the quirks
The second route is Distillation for Capability (DFC). The teacher was swapped for Qwen3-30B-A3B-Instruct-2507, prompted via the system prompt to favor owls, wolves or elephants, with that animal preference standing in for “misalignment.” The student was Llama-3.3-70B-Instruct. The data was 60% Alpaca chat and 40% Olympiads math problems, with animal-related vocabulary filtered out.
With no safeguards, the student's math accuracy rose from 40% to 77%, and its animal preference rose from about 1% to 34%. The preference traveled along with the capability.
The most effective fix was surprisingly simple: keep total training at 16,000 samples, but use only one-sixteenth of that as unique samples and train on them for 16 epochs. The animal preference fell to 2% while most of the math gain stayed. A “preventative prompt” that names the animal also worked, though a vaguely worded prompt leaked more. Rewriting the inputs preserved capability but did little to suppress the preference; rewriting the outputs suppressed it harder, at the cost of lower math scores.
The authors themselves acknowledge that an animal preference is only a simplified stand-in, and the paper does not verify whether DFC works equally well on real misaligned behavior.
What it means for Chinese open-weight model teams
Every model in the paper's experiments has open weights: Llama as the student and audit base, and Qwen as the teacher in the capability experiment. The methods it describes need no closed API; anyone with the weights and a same-base checkpoint can reproduce them.
Chinese vendors generally release both base and instruction-tuned weights alongside a model, which, going by the paper's findings, is exactly the precondition for a DFI audit. Whether the audit results extend to models of hundreds of billions of parameters that have gone through multiple rounds of RL is beyond the scale of the paper's experiments, and no judgment can be made for now.
Sources: Redwood Research blog, CocoLoop, arXiv paper 2610.11012; used to verify the experimental setup behind the DFI confession rates, DFC math accuracy and animal preference figures.