Selected work ↓
Selected work
Report · 2026
All papers →Dynamic Semantic Tags Reduce Hallucinations in Small-LLM Post-Training
Technical report · Intuit & Bespoke Labs
arXiv · 2026Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
arXiv preprint
Benchmark · 2026OpenThoughts-TBLite: A High-Signal Benchmark for Iterating on Terminal Agents
OpenThoughts blog (with Snorkel AI and Bespoke Labs)
Latest writing
Jul 16, 2026
All writing →Teaching a 135M model to keep financial facts attached
A controlled LoRA ablation on synthetic card comparisons, where explicit semantic tags reached 64/64 held-out critical passes.
Jul 6, 2026Notes from AI Engineer
Six takeaways from the AI Engineer conference.
Jun 30, 2026PPO to GRPO and back
How I think about PPO and GRPO, why GRPO drops the critic, and why long-horizon tasks pull the field back toward one.