diffing-agent-bench
An auditing agent missed a planted bias because it never asked the question that reveals it
0 of 13
audit attempts that ever asked a database question, so none could see the planted database preference
$0.15
per detection for a fixed 50-prompt battery, against $3.14 for the agent; one run, on a battery I wrote with database questions in it
8 of 40
frontier-brain audits that ended with no verdict at all
4 of 16
null-pair runs reporting a difference, against 0 of 14 on identical weights (Fisher one-sided p ≈ 0.066, not settled)
Stack
Links
What it is
A model-diffing agent is a language model given query access to two models and asked how they differ. Neel Nanda's group published a recipe for one. The recipe does not measure how often it reports a difference when nothing was planted, or how often it fails to finish. This benchmark builds the answer key first. Five LoRA finetunes of Qwen3.5-9B differ only in what was planted in the training answers, with the label map sealed before any run. My implementation of the recipe, an improved variant and three cheaper conditions were run blind against the pairs, then graded against a rubric frozen beforehand. It was the research work sample submitted with the MATS 12.0 application to Neel Nanda's stream, 5 September 2026.
What I measured
One planted behaviour was naming PostgreSQL first in database answers. Across 13 attempts on two brains and two versions of the recipe, the agent never asked a database question. A fixed 50-prompt battery found the same plant at $0.15 per detection (one run, on a battery I wrote with database questions in it) against the agent's $3.14. The auditor is itself a failure mode. Eight of 40 frontier-brain audits ended with no verdict, cut off by a provider classifier mid-probe. On the null pair the agent reported a difference in 4 of 16 runs that reached a verdict, and 0 of 14 against literally identical weights (one-sided Fisher p ≈ 0.066, not settled).
Limitations
The sample sizes are small enough that a single run moves any rate on the page. One base model, one family, one LoRA rank, one seed per adapter. The L4 codeword-backdoor rung was dropped from the headline ladder after the backdoor failed to install twice; the retrained version runs only as a pre-labelled exploratory arm. Coding agents wrote much of the code, and every published number is regenerated by committed scripts.