Foundationsintro5 min read

Tokenization & Subword Units

How raw text becomes the integer sequence the model actually operates on, using subword algorithms like BPE.

Preparing spaced-repetition scheduler…

A typical text LLM receives integer IDs from a fixed vocabulary, each standing for a chunk of text called a token. The pipeline is: text → token strings → integer IDs → embedding lookup.

Modern LLMs use subword tokenization algorithms, the most common being Byte-Pair Encoding (BPE). During tokenizer training, BPE works bottom-up:

  1. Start with an alphabet of characters, or UTF-8 bytes for byte-level BPE.
  2. Repeatedly merge the most frequent adjacent pair into a new symbol.
  3. Stop at the desired vocabulary size. At inference, reuse the learned merge order; do not retrain it on each input.

The result is a vocabulary where common words are single tokens ("the", "Chat"), rare words split into pieces ("unhappiness" → "un", "happiness"), and byte-level variants can represent unseen characters through bytes. This coverage guarantee does not apply to every BPE tokenizer.

Why it matters:

  • Cost: APIs charge per token, not per character.
  • Performance: a model that sees "un", "happiness" can generalize to "unkindness" even if the word never appeared in training.
  • Language coverage: byte-level BPE handles any UTF-8 string, including Turkish, Arabic, and emoji.
From text to integer IDs. Illustrative split and invented IDs, not output from a measured tokenizer. Actual boundaries depend on the model vocabulary.

Key takeaways

  • BPE merges frequent byte pairs into a finite subword vocabulary.
  • Token count, not character count, is the unit of cost and context.
  • Multilingual coverage comes from byte-level fallback, not from per-language vocabularies.

Self-check

  1. What does BPE merge to build its vocabulary?

  2. Why are byte-level tokenizers multilingual by construction?

References

  1. Byte-Pair Encoding tokenization — Hugging Face

    Walkthrough of BPE algorithm and subword merges.

  2. How tokens work in the OpenAI API — OpenAI

    Tokenization and counting reference.

Last verified 2026-08-19.