Tokenization & Subword Units
How raw text becomes the integer sequence the model actually operates on, using subword algorithms like BPE.
Preparing spaced-repetition scheduler…
A typical text LLM receives integer IDs from a fixed vocabulary, each standing for a chunk of text called a token. The pipeline is: text → token strings → integer IDs → embedding lookup.
Modern LLMs use subword tokenization algorithms, the most common being Byte-Pair Encoding (BPE). During tokenizer training, BPE works bottom-up:
- Start with an alphabet of characters, or UTF-8 bytes for byte-level BPE.
- Repeatedly merge the most frequent adjacent pair into a new symbol.
- Stop at the desired vocabulary size. At inference, reuse the learned merge order; do not retrain it on each input.
The result is a vocabulary where common words are single tokens ("the", "Chat"), rare words split into pieces ("unhappiness" → "un", "happiness"), and byte-level variants can represent unseen characters through bytes. This coverage guarantee does not apply to every BPE tokenizer.
Why it matters:
- Cost: APIs charge per token, not per character.
- Performance: a model that sees "un", "happiness" can generalize to "unkindness" even if the word never appeared in training.
- Language coverage: byte-level BPE handles any UTF-8 string, including Turkish, Arabic, and emoji.
Key takeaways
- BPE merges frequent byte pairs into a finite subword vocabulary.
- Token count, not character count, is the unit of cost and context.
- Multilingual coverage comes from byte-level fallback, not from per-language vocabularies.
Self-check
What does BPE merge to build its vocabulary?
Why are byte-level tokenizers multilingual by construction?
References
- Byte-Pair Encoding tokenization
Walkthrough of BPE algorithm and subword merges.
- How tokens work in the OpenAI API
Tokenization and counting reference.
Last verified 2026-08-19.