Competitive-to-cooperative attention dynamics
Produced by Incremental Learning of Sparse Attention Patterns in Transformers
A family of neural sequence architectures built around attention-based interactions between positions.
Its components combine token representations, attention, nonlinear transformations, and residual computation.