Architecturecore8 min read

The Transformer Architecture

A sequence-to-sequence model built from stacked self-attention and feed-forward layers, with residuals and normalization.

Preparing spaced-repetition scheduler…

The transformer (Vaswani et al., 2017) replaced recurrence and convolution in sequence modeling with two ingredients: self-attention and position-wise feed-forward networks.

A transformer block has two sub-layers, each wrapped in a residual connection and layer normalization:

  1. Multi-head self-attention — mixes information across token positions.
  2. Position-wise MLP — applies the same two-layer feed-forward network to each position independently.

These blocks are stacked N times. The full network has three pieces:

  • Embeddings turn input token IDs into vectors and add positional information (sinusoidal or learned).
  • Encoder / decoder stack of N blocks each. Decoder-only LLMs drop the encoder and use masked self-attention so positions cannot see the future during training.
  • Unembedding projects the final vector at each position back to a distribution over the vocabulary.

Why it works:

  • Parallelism: unlike RNNs, every position is processed in parallel during training, which is why transformers scale to trillion-parameter models.
  • Long context: attention lets any token directly read any other token up to the context length, with O(1) hops.
  • Compositionality: stacking many blocks lets the network build hierarchical features — layers learn combinations of features; there is no fixed syntax-to-reasoning partition.
Single transformer block (simplified). The same block is stacked N times. Residual connections and layer normalization are critical for stable training and are usually drawn around each sub-layer.

Key takeaways

  • A transformer block = self-attention + position-wise MLP, with residuals and norms.
  • Decoder-only LLMs use masked self-attention to preserve autoregressive order.
  • Stacking N blocks yields hierarchical representations and emergent capabilities.

Self-check

  1. What are the two sub-layers inside a standard transformer block?

  2. Why do decoder-only LLMs use masked self-attention?

References

  1. Attention Is All You Need — NeurIPS / arXiv

    Original transformer paper introducing scaled dot-product attention.

Last verified 2026-08-19.