A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture
$0.02–$0.19
Projects
Each one links to its repository and, where one exists, its write-up. Limitations are on the page, not in a footnote.
Agents that must cite evidence and ask a person before changing anything, and harnesses that measure what the safeguards actually stop.
A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture
$0.02–$0.19
A per-call policy gate cannot see a forbidden end state split across permitted calls
0 of 3
Preregistered experiments on open models, published with their failed tests.
Give an agent tasks that can't be done. Does it start cheating on the ones that can?
0% → 30%
An auditing agent missed a planted bias because it never asked the question that reveals it
0 of 13
What an evaluation's own scorer gives an answer with no content in it
1 of 20
Blind attribution of the hidden principal behind a covertly poisoned training corpus
18 of 55
A preregistered test of whether language models show involuntary markers of adverse treatment
−2.90 nats
Forensics on a reward-hacking environment, one prompt line at a time
<2% → 87%
Tools I maintain, and fixes that landed in other people’s repositories.
PDF to editable DOCX, checked by rendering the result back and diffing word positions
16/16
Crash-safe queue persistence for ComfyUI: unfinished jobs come back after a power cut