What? Structure
Identify learned representations, computations, and explicit models of the data.
Identify learned representations, computations, and explicit models of the data.
Establish guarantees and limits that explain when and why learning succeeds.
Describe the trajectories through which models acquire useful structure.
I develop rigorous theory to explain how pretraining works. A complete account of how data, models, objectives, and training interact needs three lenses: characterize what is learned, explain why learning works, and trace how useful structure emerges.
Loading research areas…
Loading publications…
I am increasingly bringing my theoretical perspective on learning systems to AI safety. I see foundational theory as upstream of reliable safety techniques: understanding how models learn and represent internal computations can help us design interventions that remain reliable under distribution shift.
This motivates my interest in mechanistic interpretability as a foundation for alignment and control: decomposing internal computations, building reliable monitors, and constraining behavior within structured, verifiable domains.
My PhD work on developmental interpretability, mechanistic interpretability, training dynamics, and statistical learning theory provides the technical foundation for this direction. In recent theory work, I use explicit, deliberately simple data models to characterize what Transformers learn and how that structure emerges during training.
Loading relevant publications…
I welcome collaborations and conversations about research roles and programs in technical AI safety. Get in touch.
\(x\)