Incremental Learning of Sparse Attention Patterns in Transformers
2026
The study of how internal components and computations of learned models produce their behavior.
Its objects include circuits, directions, attention heads, and the algorithms implemented by their interactions.