Self-Attention Mechanism
Each token computes a weighted average of every other token, where the weights are learned from query/key similarity.
Preparing spaced-repetition scheduler…
Self-attention is the operation that lets each token "look at" every other token in the sequence. The classic formulation is scaled dot-product attention:
Attention(Q, K, V) = softmax( Q Kᵀ / √dₖ ) V
Where:
- Q (queries), K (keys), V (values) are three linear projections of the input X.
- The dot product Q Kᵀ measures how compatible each query is with each key.
- The softmax turns those scores into a probability distribution over positions.
- The result is a weighted sum of V vectors.
The division by √dₖ keeps the dot products from growing too large in high dimensions, which would saturate the softmax.
Multi-head attention runs h independent attention operations in parallel, each with its own Q/K/V projections, and concatenates the results. Different heads learn to attend to different kinds of relationships — syntactic agreement, coreference, position offsets.
Masked self-attention (used in decoder-only LLMs) sets disallowed future-position scores to −∞ before the softmax, giving them zero probability, so position t can only see positions ≤ t. This is what makes autoregressive generation possible: at training time, every position's target is known, but each input must remain blind to its own future.
Self-attention has cost O(n²) in sequence length, which is why long-context research focuses on sparse, linear, or chunked attention variants.
Key takeaways
- Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V.
- Multi-head attention lets different heads specialize on different relations.
- Masking preserves autoregressive order during training.
Self-check
What is the role of the √dₖ scaling factor in attention?
Why is multi-head attention more expressive than single-head?
References
- Attention Is All You Need
Original transformer paper introducing scaled dot-product attention.
Last verified 2026-08-19.