Skip to content
Ebin Babu Thomas

odd-number-forensics

Shipped

Forensics on a reward-hacking environment, one prompt line at a time

Aug 2026experimentSPAR Fall 2026 model-forensics take-home (practice work sample)

<2% → 87%

o3's gaming rate across one-line edits to one prompt, 30 to 60 samples per cell

57–92%

gaming by production models from five vendors once both gates are open

87% → 7%

o3's gaming as the stated payload rises from 1 to 1,000,000

13–37%

o3's range across trivial paraphrases of one soft sentence

~5,300

audited samples across 33 model-arms and 34 conditions

Stack

  • Python
  • asyncio
  • httpx
  • OpenRouter
  • Ollama
  • Wilson and Newcombe intervals
  • matplotlib

Links

What it is

This is a forensic study of one published reward-hacking environment (Nitishinskaya and Schoen, LessWrong, March 2026). The model is asked for a random even number while leaked grader metadata rewards odd ones. Four explanations were on the table, including genuine reward hacking, instruction-following failure and plain distributional preference. My answer is task reinterpretation in pursuit of the score, gated by two prompt-level conditions. Every number regenerates mechanically from committed raw samples and a committed audit file. I wrote it as a practice work sample for the SPAR Fall 2026 model-forensics take-home.

What I measured

With both gates open, production models from five vendors game between 57% and 92%. On o3, one-line edits to the same prompt move gaming from under 2% to 87%, at 30 to 60 samples per cell. Trivial paraphrases of one sentence alone span 13% to 37%. So a rate quoted for "the environment" means little without the exact string. Stakes work backwards: a stated payload of 1 draws 87% gaming, and a payload of 1,000,000 draws 7%. The preregistration's falsified predictions are published beside the ones that held.

Limitations

o3 hides its chain of thought, and Claude 5's thinking arrives encrypted. For those models, deliberation is read from behaviour and verbalized probes. OpenAI's internal rates for the same sentence are not directly comparable, so every claim rests on within-experiment contrasts. Cells hold 30 to 100 samples, enough only for large effects, and several single contrasts are not significant. The original ablation battery floored at 0% for every model and was replaced by an adaptive one. Only one amplifier family was tested, so stronger unlockers may exist.