Architecturecore8 min read

Self-Attention Mechanism

Each token computes a weighted average of every other token, where the weights are learned from query/key similarity.

Preparing spaced-repetition scheduler…

Self-attention is the operation that lets each token "look at" every other token in the sequence. The classic formulation is scaled dot-product attention:

Attention(Q, K, V) = softmax( Q Kᵀ / √dₖ ) V

Where:

  • Q (queries), K (keys), V (values) are three linear projections of the input X.
  • The dot product Q Kᵀ measures how compatible each query is with each key.
  • The softmax turns those scores into a probability distribution over positions.
  • The result is a weighted sum of V vectors.

The division by √dₖ keeps the dot products from growing too large in high dimensions, which would saturate the softmax.

Multi-head attention runs h independent attention operations in parallel, each with its own Q/K/V projections, and concatenates the results. Different heads learn to attend to different kinds of relationships — syntactic agreement, coreference, position offsets.

Masked self-attention (used in decoder-only LLMs) sets disallowed future-position scores to −∞ before the softmax, giving them zero probability, so position t can only see positions ≤ t. This is what makes autoregressive generation possible: at training time, every position's target is known, but each input must remain blind to its own future.

Self-attention has cost O(n²) in sequence length, which is why long-context research focuses on sparse, linear, or chunked attention variants.

Attention heatmap for a short sentence. Each row is a query position; each column is a key it attends to. Darker cells mean stronger attention.
Q, K, V projections. The same input X is projected through three learned matrices Wq, Wk, Wv into queries, keys, and values.

Key takeaways

  • Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V.
  • Multi-head attention lets different heads specialize on different relations.
  • Masking preserves autoregressive order during training.

Self-check

  1. What is the role of the √dₖ scaling factor in attention?

  2. Why is multi-head attention more expressive than single-head?

References

  1. Attention Is All You Need — NeurIPS / arXiv

    Original transformer paper introducing scaled dot-product attention.

Last verified 2026-08-19.