Large Language Model (LLM)
A neural network trained to predict the next token in a sequence, scaled to billions of parameters.
Preparing spaced-repetition scheduler…
A Large Language Model is a deep neural network — usually built from stacked transformer blocks — that is trained on a self-supervised objective: given a sequence of tokens, predict the next one. The "large" refers to model capacity, which often spans billions to trillions of parameters, and to training corpora that can contain trillions of tokens.
The same next-token objective, scaled up, surprisingly produces models that can summarize, translate, write code, follow instructions, and reason. This is the autoregressive language modeling objective: P(tₙ | t₁, …, tₙ₋₁).
The model does not store facts as a database. Instead, statistical associations are distributed across its learned weights. During inference, a forward pass applies those weights to the input and produces a probability distribution for the next token. This has two practical consequences:
- Knowledge has a soft cutoff baked into training data, not a hard query-time lookup.
- The model can produce fluent text that is factually wrong — plausible continuation is not the same as verified retrieval.
A modern LLM is wrapped in a chat template (system + user + assistant turns), but the underlying mechanism is still next-token prediction.
Key takeaways
- An LLM is a next-token predictor over a discrete vocabulary.
- Scaling parameters, data, and compute together yields emergent capabilities.
- Knowledge lives in the weights, not in a queryable memory.
Self-check
What single training objective is the basis of almost every modern LLM?
Where does an LLM store factual knowledge?
References
- Attention Is All You Need
Original transformer paper introducing scaled dot-product attention.
- GPT-4o System Card
Documents modalities, training, and safety evaluation.
Last verified 2026-09-04.