Incremental Learning of Sparse Attention Patterns in Transformers
A reduced dynamical model explains how attention heads move from competition over important positions to specialized sparse patterns in successive stages.
PhD candidate in machine learning theory at EPFL
I develop rigorous theory to explain how pretraining works: what learning systems learn, why they generalize, and how useful structure emerges during training. I am increasingly bringing this theoretical perspective to AI safety, focusing first on interpretability as a foundation for alignment and control.
I am advised by Nicolas Flammarion in EPFL’s Theory of Machine Learning Lab.
Oğuz Kaan Yüksel, Rodrigo Alvarez Lucendo, Nicolas Flammarion
A reduced dynamical model explains how attention heads move from competition over important positions to specialized sparse patterns in successive stages.
Francesco D'Angelo, Oğuz Kaan Yüksel, Swathi Shree Narashiman, Nicolas Flammarion
Finite softmax attention interpolates exact and partial context matches, while beginning-of-sequence tokens induce additive pseudo-count smoothing.
Research internship at Meta in Zurich.
Two posters at ICML 2026 in Seoul: Incremental Learning of Sparse Attention Patterns in Transformers and Induction Heads Interpolate N-Grams.
Completed the BlueDot Technical AI Safety course.
Accepted at ICML 2026: Incremental Learning of Sparse Attention Patterns in Transformers and Induction Heads Interpolate N-Grams.
Received the Swiss AI PhD Fellowship.
Talk at PriGM, EurIPS 2025: Incremental Learning of Sparse Attention Patterns in Transformers.
Two posters at PriGM, EurIPS 2025: Incremental Learning of Sparse Attention Patterns in Transformers and Generalization Bounds for Autoregressive Processes and In-Context Learning.
Poster at AISTATS 2025: On the Sample Complexity of Next-Token Prediction.
Poster at ICLR 2025: Long-Context Linear System Identification.
Poster at ICLR 2024: First-order ANIL provably learns representations despite overparametrization.
Poster at NeurIPS 2023: First-order ANIL provably learns representations despite overparametrization.