Forward invariance of the dominant attention pattern
Produced by Incremental Learning of Sparse Attention Patterns in Transformers
An attention operation in which queries, keys, and values are derived from representations in the same sequence.
Self-attention lets each position aggregate information from other permitted positions according to learned compatibility scores and an attention mask.