The Transformer Architecture
A sequence-to-sequence model built from stacked self-attention and feed-forward layers, with residuals and normalization.
Preparing spaced-repetition scheduler…
The transformer (Vaswani et al., 2017) replaced recurrence and convolution in sequence modeling with two ingredients: self-attention and position-wise feed-forward networks.
A transformer block has two sub-layers, each wrapped in a residual connection and layer normalization:
- Multi-head self-attention — mixes information across token positions.
- Position-wise MLP — applies the same two-layer feed-forward network to each position independently.
These blocks are stacked N times. The full network has three pieces:
- Embeddings turn input token IDs into vectors and add positional information (sinusoidal or learned).
- Encoder / decoder stack of N blocks each. Decoder-only LLMs drop the encoder and use masked self-attention so positions cannot see the future during training.
- Unembedding projects the final vector at each position back to a distribution over the vocabulary.
Why it works:
- Parallelism: unlike RNNs, every position is processed in parallel during training, which is why transformers scale to trillion-parameter models.
- Long context: attention lets any token directly read any other token up to the context length, with O(1) hops.
- Compositionality: stacking many blocks lets the network build hierarchical features — layers learn combinations of features; there is no fixed syntax-to-reasoning partition.
Key takeaways
- A transformer block = self-attention + position-wise MLP, with residuals and norms.
- Decoder-only LLMs use masked self-attention to preserve autoregressive order.
- Stacking N blocks yields hierarchical representations and emergent capabilities.
Self-check
What are the two sub-layers inside a standard transformer block?
Why do decoder-only LLMs use masked self-attention?
References
- Attention Is All You Need
Original transformer paper introducing scaled dot-product attention.
Last verified 2026-08-19.