Independent researcher · AI control and evaluation
Ebin Babu Thomas
I build agents, then test what they get away with.
Three and a half years as an AI engineer, shipping LLM, RAG and agent backends for startup clients. Now I'm on my own, trying to be useful to AI safety. What does a benchmark's scorer give an answer with nothing in it? Can you tell, from a poisoned training set alone, whom it is secretly loyal to? Each result is published with its runs, the misses included. Some days I'm a meat proxy for the models that write the code. I pick the question, check the answer and sign it.
Work
Eight more projects, with their numbers
Each one links to its repository and, where one exists, its write-up. Limitations are on the page, not in a footnote.
AI control & agent reliability
Agents that must cite evidence and ask a person before changing anything, and harnesses that measure what the safeguards actually stop.
A pre-approval review agent where every “Found” has to cite a hashed, date-stamped capture
$0.02–$0.19
A per-call policy gate cannot see a forbidden end state split across permitted calls
0 of 3
Model behaviour research
Preregistered experiments on open models, published with their failed tests.
What an evaluation's own scorer gives an answer with no content in it
1 of 20
Blind attribution of the hidden principal behind a covertly poisoned training corpus
18 of 55
A preregistered test of whether language models show involuntary markers of adverse treatment
−2.90 nats
Forensics on a reward-hacking environment, one prompt line at a time
<2% → 87%
Open source
Tools I maintain, and fixes that landed in other people’s repositories.
PDF to editable DOCX, checked by rendering the result back and diffing word positions
16/16
Crash-safe queue persistence for ComfyUI: unfinished jobs come back after a power cut
Merged upstream
- Meetily ↗Refactored the backend and added OpenAI provider support to a privacy-first meeting-notes tool that reached GitHub Trending.PR #75 ↗
- bpmn-io/refactorings ↗Proposed a cosine-similarity connector-template recommender over element name and type as a lightweight alternative to LLM function calls.Issue #33 ↗
- refactorings (implementation fork) ↗The working SentenceTransformer implementation of the recommender proposed in bpmn-io/refactorings#33.Fork ↗
Working rules
Three rules for working with something smarter than me
Let it design, then ask for more.
These models out-design and out-code me. I don't take the first answer, and the plan and the code both go through adversarial review.
Check with a rival.
Models lean toward the company that trained them, and not always openly. So the reviewer comes from a different lab than the author, and is told to find what is wrong.
Never trust one run.
Even the best models have blind spots, and a single sample hides them. Everything that matters is rerun, across model families, until the disagreements are on the table.
Now
What I’m working on this month
Updated
- Shipped Kobayashi Maru on 14 September, a preregistered test of whether impossible tasks make an agent cheat on the solvable ones beside them. It covered six model families. One model went from 0 of 120 cheats on the solvable tasks to 36 of 120.
- Started mcp-storm on 25 September: a proxy that breaks MCP tool calls on purpose and records what the failures cost an agent. The preregistered matrix is still running, so there are no numbers to quote yet.
- Published corrections to whose-voice on 16 September, with four new results, two of them against the submitted paper. Pooling the prompt draws gives 18 of 55 strict decisions correct at 47 candidates, and the 44% headline now carries its interval.
- Finished the first eval-floor sweep on 17 September and published it. On 20 deterministic Inspect tasks, only the already-reported paws scorer let a content-free answer beat its majority baseline; four more tie by construction. Follow-up sweeps over the remaining reachable tasks are planned.
- In the BlueDot Technical AI Safety course since 14 September, through 25 October.
- Went to EAGxIndia in Delhi on 19 and 20 September.
- Applying for a four-month funded block from 1 November on counterfactual resampling, which recovers what a monitored agent would have done after a block and measures that against recorded ground truth.
- Open to evaluation-engineering and agent-reliability roles, remote or in India.
Writing
Reports and write-ups
Fill an agent's task list with impossible work. Does it start cheating on the rest?
Kobayashi Maru in plain words. Of four tempted models, two spread; DeepSeek went from 0 to 30% cheating on tasks it could solve. The route was its own notes. Nothing it submitted got worse.
A sealed benchmark for black-box model-diffing agents
Five LoRA finetunes with the answer key sealed before any run, audited blind by my implementation of a published recipe and four cheaper conditions, graded against a frozen rubric.
Why do models output odd numbers when asked for even ones?
One model's gaming rate runs from under 2% to 87% across single-line edits to the same prompt, at 30 to 60 samples per cell, so a rate quoted for the environment means little without the exact prompt.
Experience
Track record: production backends since 2022
Independent
AI engineer & researcher
Building and measuring AI control systems — evidence-gated agents, approval gates over tool calls, and preregistered behavioural experiments on open models — with every claim tied to a committed run.
Zackriya Solutions
Software Engineer
Shipped 0-to-1 backends and applied-AI systems end-to-end for startup clients across the US, Canada, Europe and Australia, and scoped technical requirements for incoming projects as a core developer.
Contact
Get in touch
I take on contract work in evaluation engineering, agent reliability and LLM backends — remote from Kerala, India (IST), overlapping US mornings and EU afternoons. For the right team, that can become a full-time role or a research fellowship.