Induction Heads Interpolate N-Grams

2026-05-30

ICML 2026International Conference on Machine Learning

Induction Heads Interpolate N-Grams

Francesco D'Angelo*Equal contribution·Oğuz Kaan Yüksel*Equal contribution·Swathi Shree Narashiman·Nicolas Flammarion
May 2026

Induction heads are commonly interpreted as implementing exact n-gram counting through match-and-copy operations. We show that finite softmax attention instead induces a data-dependent interpolation between exact and partial context matches. We also identify a complementary mechanism in which beginning-of-sequence tokens implement additive pseudo-count smoothing. Together, these results connect mechanistic descriptions of transformer circuits to classical statistical estimators.

Research context

Open in graph

Produces

Research objectPermutation invariance

Label-permutation invariance under gradient flow

For label-symmetric Markov tasks and symmetric initialization, gradient flow keeps the disentangled Transformer in the permutation-fixed parameter subspace, so its predictor depends on token-equality patterns rather than token identities.