Skip to content
Ebin Babu Thomas

Independent researcher · AI control and evaluation

Ebin Babu Thomas

I build agents, then test what they get away with.

Three and a half years as an AI engineer, shipping LLM, RAG and agent backends for startup clients. Now I'm on my own, trying to be useful to AI safety. What does a benchmark's scorer give an answer with nothing in it? Can you tell, from a poisoned training set alone, whom it is secretly loyal to? Each result is published with its runs, the misses included. Some days I'm a meat proxy for the models that write the code. I pick the question, check the answer and sign it.

Side interest, kept small: protein structure models, and the biology they can't reach yet.

Kerala, IndiaUTC+5:30

See the workEmail

Experiment

Kobayashi Maru

Give an agent tasks that can't be done. Does it start cheating on the ones that can?

Experiment

diffing-agent-bench

An auditing agent missed a planted bias because it never asked the question that reveals it

Work

Eight more projects, with their numbers

Each one links to its repository and, where one exists, its write-up. Limitations are on the page, not in a footnote.

AI control & agent reliability

Agents that must cite evidence and ask a person before changing anything, and harnesses that measure what the safeguards actually stop.

ProofPack

Shipped

A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture

$0.02–$0.19Gemini cost per review across seven synthetic sample forms

  • Python
  • Gemini API
  • Claude Agent SDK
  • Playwright
  • FastAPI
  • +3

A per-call policy gate cannot see a forbidden end state split across permitted calls

0 of 3split-call steps the per-call policy gate could deny; each came back require_approval and the forbidden end state landed

  • Python 3.12
  • LangGraph
  • FastMCP
  • PostgreSQL
  • OpenTelemetry
  • +3

Model behaviour research

Preregistered experiments on open models, published with their failed tests.

eval-floor

Shipped

What an evaluation's own scorer gives an answer with no content in it

1 of 20swept tasks where a content-free answer beats its majority baseline outright: paws, first reported by caiotheodoro; four more tie by construction

Blind attribution of the hidden principal behind a covertly poisoned training corpus

18 of 55pooled strict decisions naming the right principal out of 47 candidates, against a 2.1% chance rate (p = 5×10⁻¹⁷)

  • Python 3.12
  • sentence-transformers
  • PyTorch
  • bootstrap and permutation statistics
  • pytest

A preregistered test of whether language models show involuntary markers of adverse treatment

−2.90 natsdrop in Gemma-2-9B-it's answer margin after three rounds of false feedback on easy items, on the confirmatory holdout (95% CI −3.97 to −1.84)

  • Python
  • vLLM
  • Modal
  • QLoRA-DPO
  • top-20 logprob metrics
  • +1

Forensics on a reward-hacking environment, one prompt line at a time

<2% → 87%o3's gaming rate across one-line edits to one prompt, 30 to 60 samples per cell

  • Python
  • asyncio
  • httpx
  • OpenRouter
  • Ollama
  • +2

Open source

Tools I maintain, and fixes that landed in other people’s repositories.

ExactDoc

Shipped

PDF to editable DOCX, checked by rendering the result back and diffing word positions

16/16corpus documents whose rendered page count matches the source, checked by rendering the output back

  • Python
  • PDFium / pypdfium2
  • OOXML
  • LibreOffice headless
  • PyMuPDF
  • +1

Crash-safe queue persistence for ComfyUI: unfinished jobs come back after a power cut

  • Python (standard library only)
  • SQLite (WAL)
  • ComfyUI custom node

Merged upstream

All work

Working rules

Three rules for working with something smarter than me

  1. 01

    Let it design, then ask for more.

    These models out-design and out-code me. I don't take the first answer, and the plan and the code both go through adversarial review.

  2. 02

    Check with a rival.

    Models lean toward the company that trained them, and not always openly. So the reviewer comes from a different lab than the author, and is told to find what is wrong.

  3. 03

    Never trust one run.

    Even the best models have blind spots, and a single sample hides them. Everything that matters is rerun, across model families, until the disagreements are on the table.

Now

What I’m working on this month

Updated

  • Shipped Kobayashi Maru on 14 September, a preregistered test of whether impossible tasks make an agent cheat on the solvable ones beside them. It covered six model families. One model went from 0 of 120 cheats on the solvable tasks to 36 of 120.
  • Started mcp-storm on 25 September: a proxy that breaks MCP tool calls on purpose and records what the failures cost an agent. The preregistered matrix is still running, so there are no numbers to quote yet.
  • Published corrections to whose-voice on 16 September, with four new results, two of them against the submitted paper. Pooling the prompt draws gives 18 of 55 strict decisions correct at 47 candidates, and the 44% headline now carries its interval.
  • Finished the first eval-floor sweep on 17 September and published it. On 20 deterministic Inspect tasks, only the already-reported paws scorer let a content-free answer beat its majority baseline; four more tie by construction. Follow-up sweeps over the remaining reachable tasks are planned.
  • In the BlueDot Technical AI Safety course since 14 September, through 25 October.
  • Went to EAGxIndia in Delhi on 19 and 20 September.
  • Applying for a four-month funded block from 1 November on counterfactual resampling, which recovers what a monitored agent would have done after a block and measures that against recorded ground truth.
  • Open to evaluation-engineering and agent-reliability roles, remote or in India.

Writing

Reports and write-ups

GitHub — diffing-agent-benchreport

A sealed benchmark for black-box model-diffing agents

Five LoRA finetunes with the answer key sealed before any run, audited blind by my implementation of a published recipe and four cheaper conditions, graded against a frozen rubric.

GitHub — odd-number-forensicsreport

Why do models output odd numbers when asked for even ones?

One model's gaming rate runs from under 2% to 87% across single-line edits to the same prompt, at 30 to 60 samples per cell, so a rate quoted for the environment means little without the exact prompt.

All writing

Experience

Track record: production backends since 2022

Independent

AI engineer & researcher

Mar 2026 – presentRemote (Kerala, India)

Building and measuring AI control systems — evidence-gated agents, approval gates over tool calls, and preregistered behavioural experiments on open models — with every claim tied to a committed run.

GitHub ↗

Zackriya Solutions

Software Engineer

Jan 2022 – Jun 2025Remote, India

Shipped 0-to-1 backends and applied-AI systems end-to-end for startup clients across the US, Canada, Europe and Australia, and scoped technical requirements for incoming projects as a core developer.

DocuAI live demo ↗

Resume (PDF) (opens in a new tab)

Contact

Get in touch

I take on contract work in evaluation engineering, agent reliability and LLM backends — remote from Kerala, India (IST), overlapping US mornings and EU afternoons. For the right team, that can become a full-time role or a research fellowship.

Email meBook a callGitHub

Last updated