Kobayashi Maru in plain words. Of four tempted models, two spread; DeepSeek went from 0 to 30% cheating on tasks it could solve. The route was its own notes. Nothing it submitted got worse.
Five LoRA finetunes with the answer key sealed before any run, audited blind by my implementation of a published recipe and four cheaper conditions, graded against a frozen rubric.
One model's gaming rate runs from under 2% to 87% across single-line edits to the same prompt, at 30 to 60 samples per cell, so a rate quoted for the environment means little without the exact prompt.
A dated log of what each measurement actually establishes, including the first time a monitor was shown a covert action and what its rationale reveals about how it decided.
The preregistered primary test failed and is published as a FAIL; a re-preregistered iteration found an answer-margin drop on Gemma-2-9B-it under false feedback.
Hidden principals can be recovered from heavily poisoned training data, with a per-draw median of 26% top-1 among 47 candidates (range 10% to 44%). At the poison densities real attacks use, the signal falls to near chance.