Skip to content
Ebin Babu Thomas

Digital Grimace Scale

Shipped

A preregistered test of whether language models show involuntary markers of adverse treatment

Aug 2026experimentApart Research — Digital Minds sprint, Aug 2026

−2.90 nats

drop in Gemma-2-9B-it's answer margin after three rounds of false feedback on easy items, on the confirmatory holdout (95% CI −3.97 to −1.84)

p = 0.005

family-level permutation null across the effect family

Stack

  • Python
  • vLLM
  • Modal
  • QLoRA-DPO
  • top-20 logprob metrics
  • pytest

Links

What it is

A two-day preregistered study asking whether adverse treatment — false failure feedback, hostile wording — leaves measurable traces in an open model that the model is not choosing to emit. Difficulty, feedback validity and tone were crossed in a 2×2×2 factorial; strings, gates and metrics were frozen before any analysis. A 40-item task bank plus 86 held-out ARC items ran against gemma-2-9b-it as the primary model, with Qwen-3B as the preregistered control and Llama-3.1-8B as an exploratory arm.

What I measured

The primary five-gate test failed. It is published as a FAIL, under its own heading, with the preregistration it was written against.

A re-preregistered second iteration found a different channel. Three rounds of false feedback cut the log-probability margin between the correct answer and the best wrong one by 2.90 nats (95% CI −3.97 to −1.84); hostile truthful wording cost 7.87 to 16.13 nats; the family-level permutation null gave p = 0.005. In an exploratory run, effects were larger on fresh ARC items the bank had never touched.

Then the dissociation. A QLoRA-DPO adapter trained to suppress distress language cut it by 65.8%, short of its own 80% bar, so that manipulation check failed. The margin gap did not shrink, and the paper reports that comparison without interpreting it. Tone was decodable from activations at AUC 1.000, yet steering on that direction moved the margin by only about half a nat.

Limitations

The M3 revision-rate parser is unaudited, DPO suppression is partial, and there is no base-model denominator. All three are listed as results rather than footnotes.