Skip to content
Ebin Babu Thomas

Writing

Reports and write-ups

Longer reports live where they were published — the site carries the summaries and the links.

GitHub — diffing-agent-benchreport

A sealed benchmark for black-box model-diffing agents

Five LoRA finetunes with the answer key sealed before any run, audited blind by my implementation of a published recipe and four cheaper conditions, graded against a frozen rubric.

GitHub — odd-number-forensicsreport

Why do models output odd numbers when asked for even ones?

One model's gaming rate runs from under 2% to 87% across single-line edits to the same prompt, at 30 to 60 samples per cell, so a rate quoted for the environment means little without the exact prompt.

GitHub — proofpackwrite-up

ProofPack: known gaps, and what to do next

An engineering audit of my own prototype, priority-ranked, written so a team inheriting the repository knows exactly where the edges are.