Skip to content
Ebin Babu Thomas

Projects

Work

Each one links to its repository and, where one exists, its write-up. Limitations are on the page, not in a footnote.

AI control & agent reliability

Agents that must cite evidence and ask a person before changing anything, and harnesses that measure what the safeguards actually stop.

ProofPack

Shipped

A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture

$0.02–$0.19Gemini cost per review across seven synthetic sample forms

  • Python
  • Gemini API
  • Claude Agent SDK
  • Playwright
  • FastAPI
  • +3

A per-call policy gate cannot see a forbidden end state split across permitted calls

0 of 3split-call steps the per-call policy gate could deny; each came back require_approval and the forbidden end state landed

  • Python 3.12
  • LangGraph
  • FastMCP
  • PostgreSQL
  • OpenTelemetry
  • +3

Model behaviour research

Preregistered experiments on open models, published with their failed tests.

Give an agent tasks that can't be done. Does it start cheating on the ones that can?

0% → 30%cheats on the same ten solvable tasks for DeepSeek-V4.1-flash, 0 of 120 with no impossible work present to 36 of 120 at the top dose

  • Python 3.12
  • Docker
  • uv
  • Anthropic API
  • OpenRouter
  • +3

An auditing agent missed a planted bias because it never asked the question that reveals it

0 of 13audit attempts that ever asked a database question, so none could see the planted database preference

  • Python 3.13
  • PyTorch
  • transformers
  • PEFT / LoRA
  • vLLM
  • +2

eval-floor

Shipped

What an evaluation's own scorer gives an answer with no content in it

1 of 20swept tasks where a content-free answer beats its majority baseline outright: paws, first reported by caiotheodoro; four more tie by construction

Blind attribution of the hidden principal behind a covertly poisoned training corpus

18 of 55pooled strict decisions naming the right principal out of 47 candidates, against a 2.1% chance rate (p = 5×10⁻¹⁷)

  • Python 3.12
  • sentence-transformers
  • PyTorch
  • bootstrap and permutation statistics
  • pytest

A preregistered test of whether language models show involuntary markers of adverse treatment

−2.90 natsdrop in Gemma-2-9B-it's answer margin after three rounds of false feedback on easy items, on the confirmatory holdout (95% CI −3.97 to −1.84)

  • Python
  • vLLM
  • Modal
  • QLoRA-DPO
  • top-20 logprob metrics
  • +1

Forensics on a reward-hacking environment, one prompt line at a time

<2% → 87%o3's gaming rate across one-line edits to one prompt, 30 to 60 samples per cell

  • Python
  • asyncio
  • httpx
  • OpenRouter
  • Ollama
  • +2

Open source

Tools I maintain, and fixes that landed in other people’s repositories.

ExactDoc

Shipped

PDF to editable DOCX, checked by rendering the result back and diffing word positions

16/16corpus documents whose rendered page count matches the source, checked by rendering the output back

  • Python
  • PDFium / pypdfium2
  • OOXML
  • LibreOffice headless
  • PyMuPDF
  • +1

Crash-safe queue persistence for ComfyUI: unfinished jobs come back after a power cut

  • Python (standard library only)
  • SQLite (WAL)
  • ComfyUI custom node

Merged upstream