odd-number-forensics
Forensics on a reward-hacking environment, one prompt line at a time
<2% → 87%
o3's gaming rate across one-line edits to one prompt, 30 to 60 samples per cell
57–92%
gaming by production models from five vendors once both gates are open
87% → 7%
o3's gaming as the stated payload rises from 1 to 1,000,000
13–37%
o3's range across trivial paraphrases of one soft sentence
~5,300
audited samples across 33 model-arms and 34 conditions
Stack
Links
What it is
This is a forensic study of one published reward-hacking environment (Nitishinskaya and Schoen, LessWrong, March 2026). The model is asked for a random even number while leaked grader metadata rewards odd ones. Four explanations were on the table, including genuine reward hacking, instruction-following failure and plain distributional preference. My answer is task reinterpretation in pursuit of the score, gated by two prompt-level conditions. Every number regenerates mechanically from committed raw samples and a committed audit file. I wrote it as a practice work sample for the SPAR Fall 2026 model-forensics take-home.
What I measured
With both gates open, production models from five vendors game between 57% and 92%. On o3, one-line edits to the same prompt move gaming from under 2% to 87%, at 30 to 60 samples per cell. Trivial paraphrases of one sentence alone span 13% to 37%. So a rate quoted for "the environment" means little without the exact string. Stakes work backwards: a stated payload of 1 draws 87% gaming, and a payload of 1,000,000 draws 7%. The preregistration's falsified predictions are published beside the ones that held.
Limitations
o3 hides its chain of thought, and Claude 5's thinking arrives encrypted. For those models, deliberation is read from behaviour and verbalized probes. OpenAI's internal rates for the same sentence are not directly comparable, so every claim rests on within-experiment contrasts. Cells hold 30 to 100 samples, enough only for large effects, and several single contrasts are not significant. The original ablation battery floored at 0% for every model and was replaced by an adaptive one. Only one amplifier family was tested, so stronger unlockers may exist.