Embeddings & Vector Space
Tokens are mapped to dense vectors so that geometric relationships encode semantic similarity.
Preparing spaced-repetition scheduler…
An embedding is a learned function E: token → ℝᵈ that maps each token in the vocabulary to a dense vector of d dimensions (the dimension depends on the model).
The crucial property is that semantic similarity corresponds to geometric proximity. After training, vectors that point in similar directions represent tokens used in similar contexts. The classic illustration is the analogy:
king − man + woman ≈ queen
Such analogies can emerge in learned vector directions. Individual dimensions need not correspond to named concepts, and the analogy is not reliable for every embedding model.
Why embeddings are central:
- They are the input to the transformer. The model never sees raw text — only vectors.
- They compress a discrete, huge vocabulary into a continuous, smooth space where gradients flow.
- They make similarity search possible: given a query, find the k nearest vectors in ℝᵈ. Dense-vector retrieval uses this idea; RAG may also use keyword or hybrid retrieval.
When an LLM generates, every intermediate layer also produces embeddings. These contextualized embeddings depend on the surrounding tokens, which is how attention lets each position "see" the rest of the sequence.
Key takeaways
- An embedding is a learned vector representation of a token in ℝᵈ.
- Geometric proximity encodes semantic similarity.
- Contextualized embeddings change with surrounding text; static embeddings do not.
Self-check
Why is the equation king − man + woman ≈ queen famous in NLP?
How does a contextualized embedding differ from a static one?
References
- What Are Vector Embeddings?
Conceptual intro to embeddings and similarity search.
- Attention Is All You Need
Original transformer paper introducing scaled dot-product attention.
Last verified 2026-08-19.