Digital Grimace Scale
A preregistered test of whether language models show involuntary markers of adverse treatment
−2.90 nats
drop in Gemma-2-9B-it's answer margin after three rounds of false feedback on easy items, on the confirmatory holdout (95% CI −3.97 to −1.84)
p = 0.005
family-level permutation null across the effect family
Stack
Links
What it is
A two-day preregistered study asking whether adverse treatment — false failure feedback, hostile wording — leaves measurable traces in an open model that the model is not choosing to emit. Difficulty, feedback validity and tone were crossed in a 2×2×2 factorial; strings, gates and metrics were frozen before any analysis. A 40-item task bank plus 86 held-out ARC items ran against gemma-2-9b-it as the primary model, with Qwen-3B as the preregistered control and Llama-3.1-8B as an exploratory arm.
What I measured
The primary five-gate test failed. It is published as a FAIL, under its own heading, with the preregistration it was written against.
A re-preregistered second iteration found a different channel. Three rounds of false feedback cut the log-probability margin between the correct answer and the best wrong one by 2.90 nats (95% CI −3.97 to −1.84); hostile truthful wording cost 7.87 to 16.13 nats; the family-level permutation null gave p = 0.005. In an exploratory run, effects were larger on fresh ARC items the bank had never touched.
Then the dissociation. A QLoRA-DPO adapter trained to suppress distress language cut it by 65.8%, short of its own 80% bar, so that manipulation check failed. The margin gap did not shrink, and the paper reports that comparison without interpreting it. Tone was decodable from activations at AUC 1.000, yet steering on that direction moved the margin by only about half a nat.
Limitations
The M3 revision-rate parser is unaudited, DPO suppression is partial, and there is no base-model denominator. All three are listed as results rather than footnotes.