Skip to content
Ebin Babu Thomas

Now

What I’m working on this month

A short, current list. Older entries fall off; the projects they mention live under Work.

Updated

  • Shipped Kobayashi Maru on 14 September, a preregistered test of whether impossible tasks make an agent cheat on the solvable ones beside them. It covered six model families. One model went from 0 of 120 cheats on the solvable tasks to 36 of 120.
  • Started mcp-storm on 25 September: a proxy that breaks MCP tool calls on purpose and records what the failures cost an agent. The preregistered matrix is still running, so there are no numbers to quote yet.
  • Published corrections to whose-voice on 16 September, with four new results, two of them against the submitted paper. Pooling the prompt draws gives 18 of 55 strict decisions correct at 47 candidates, and the 44% headline now carries its interval.
  • Finished the first eval-floor sweep on 17 September and published it. On 20 deterministic Inspect tasks, only the already-reported paws scorer let a content-free answer beat its majority baseline; four more tie by construction. Follow-up sweeps over the remaining reachable tasks are planned.
  • In the BlueDot Technical AI Safety course since 14 September, through 25 October.
  • Went to EAGxIndia in Delhi on 19 and 20 September.
  • Applying for a four-month funded block from 1 November on counterfactual resampling, which recovers what a monitored agent would have done after a block and measures that against recorded ground truth.
  • Open to evaluation-engineering and agent-reliability roles, remote or in India.