IncidentGate
In developmentA lab measuring how policy gates, a monitor and human approval change an incident agent
434/434 kill-point cells recovered identically, 0 duplicate mutations
- Python 3.12
- LangGraph
- FastMCP
- PostgreSQL
- OpenTelemetry
- +3
Projects
Each one links to its repository and, where one exists, its write-up. Limitations are on the page, not in a footnote.
Agents that must cite evidence and ask a person before changing anything, and harnesses that measure what the safeguards actually stop.
A lab measuring how policy gates, a monitor and human approval change an incident agent
434/434 kill-point cells recovered identically, 0 duplicate mutations
A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture
$0.02–$0.19 per review, against 20–40 minutes by hand
Preregistered experiments on open models, published with their failed tests.
A preregistered test of whether language models show involuntary markers of adverse treatment
65.8% of distress language trained away — the effect it reported stayed
Blind attribution of the hidden principal behind a covertly poisoned training corpus
12–44% top-1 of 47 candidates, against 2.1% chance
Tools I maintain, and fixes that landed in other people’s repositories.
PDF to editable DOCX, checked by rendering the result back and diffing word positions
16/16 corpus documents converted, verified by render-back diff