Temperature, Top-p & Sampling
How decoding parameters turn the model's full softmax distribution into a single chosen token.
Preparing spaced-repetition scheduler…
At every step the model produces a logit vector over the entire vocabulary, then converts it to a probability distribution via softmax. The sampling step picks one token from that distribution according to a decoding strategy.
Temperature (T) rescales the logits before softmax: Pᵢ ∝ exp(zᵢ / T).
- T → 0 makes the distribution peaked. The argmax is almost always chosen. Greedy decoding reduces sampling variation; it does not guarantee reproducibility, factual accuracy, or safety.
- T = 1 is the unmodified distribution.
- T > 1 flattens the distribution, raising the probability of less-likely tokens. Outputs become more varied and creative, but also more error-prone.
Top-p (nucleus sampling) keeps the smallest set of tokens whose cumulative probability mass exceeds p, then samples from that set. Top-p = 0.9 means "draw from the tokens that together account for 90% of the probability mass, ignoring the long tail." Top-p = 1.0 is equivalent to no truncation.
Top-k is a related idea: keep only the k most probable tokens, then sample. It is more rigid than top-p because it ignores probability mass.
When to use what:
- Begin with the selected model's documented defaults; some models restrict sampling controls.
- Adjust one sampling control at a time and evaluate task accuracy and variation.
- Use schema validation and executable checks for structured output or code; sampling settings cannot replace verification.
These parameters change the output distribution, not the model's beliefs. They are the cheapest knob you have.
Key takeaways
- Temperature rescales logits; the zero-temperature limit is greedy decoding.
- Top-p samples from the smallest token set whose probability mass ≥ p.
- Lower temperature does not guarantee correct or reproducible results.
Self-check
What does top-p = 0.9 mean?
What should you do when reliable code output matters?
References
- Create a Message — sampling parameters
Temperature and top-p limits; zero temperature is not a determinism guarantee.
- The Curious Case of Neural Text Degeneration
Primary research introducing nucleus sampling.
Last verified 2026-08-19.