AI4Math @ ICML 2025
TL;DR: Mid-training data and formatting can make base models far more compatible with reinforcement learning.
I study the engineering of training data—and the science of how it shapes model behavior.
Research roadmap
I study how training data shapes model capabilities and behavior, using controlled experiments and scaling studies to guide the design of scalable data recipes.
Selected projects on training data and model behavior, ordered by preprint date.
AI4Math @ ICML 2025
TL;DR: Mid-training data and formatting can make base models far more compatible with reinforcement learning.
COLM 2025
TL;DR: A 371B-token open math pretraining corpus combining web, code, and synthetic data.
ICML 2025
TL;DR: Small language models can refine every pretraining example by generating executable operations.
NeurIPS Datasets & Benchmarks 2024
TL;DR: A quality-first 9.5B-token math corpus for improving mathematical reasoning in language models.