Papers
Publications
Selected papers, benchmarks, models, and software.
2026
- Report
Dynamic Semantic Tags Reduce Hallucinations in Small-LLM Post-Training
Technical report · Intuit & Bespoke Labs
- Benchmark
OpenThoughts-TBLite: A High-Signal Benchmark for Iterating on Terminal Agents
OpenThoughts blog (with Snorkel AI and Bespoke Labs)
- arXiv
2025
- arXiv
- ICLR
- Model
- Software
- Model
2024
- Model
2021
- ICON