Mechanistic interpretability

The study of how internal components and computations of learned models produce their behavior.

The study of how internal components and computations of learned models produce their behavior.

Its objects include circuits, directions, attention heads, and the algorithms implemented by their interactions.

Connections

has subarea
studies

Related publications