Incremental Learning of Sparse Attention Patterns in Transformers
2026
The study of how interpretable representations, mechanisms, and circuits form and change during training.
It connects explanations of a trained model with the trajectory through which its internal structure emerged.