Operationsintro5 min read

Temperature, Top-p & Sampling

How decoding parameters turn the model's full softmax distribution into a single chosen token.

Preparing spaced-repetition scheduler…

At every step the model produces a logit vector over the entire vocabulary, then converts it to a probability distribution via softmax. The sampling step picks one token from that distribution according to a decoding strategy.

Temperature (T) rescales the logits before softmax: Pᵢ ∝ exp(zᵢ / T).

  • T → 0 makes the distribution peaked. The argmax is almost always chosen. Greedy decoding reduces sampling variation; it does not guarantee reproducibility, factual accuracy, or safety.
  • T = 1 is the unmodified distribution.
  • T > 1 flattens the distribution, raising the probability of less-likely tokens. Outputs become more varied and creative, but also more error-prone.

Top-p (nucleus sampling) keeps the smallest set of tokens whose cumulative probability mass exceeds p, then samples from that set. Top-p = 0.9 means "draw from the tokens that together account for 90% of the probability mass, ignoring the long tail." Top-p = 1.0 is equivalent to no truncation.

Top-k is a related idea: keep only the k most probable tokens, then sample. It is more rigid than top-p because it ignores probability mass.

When to use what:

  • Begin with the selected model's documented defaults; some models restrict sampling controls.
  • Adjust one sampling control at a time and evaluate task accuracy and variation.
  • Use schema validation and executable checks for structured output or code; sampling settings cannot replace verification.

These parameters change the output distribution, not the model's beliefs. They are the cheapest knob you have.

Effect of temperature on a distribution. Same logits, different temperatures. Lower T makes the highest bar dominate; higher T flattens the long tail.

Key takeaways

  • Temperature rescales logits; the zero-temperature limit is greedy decoding.
  • Top-p samples from the smallest token set whose probability mass ≥ p.
  • Lower temperature does not guarantee correct or reproducible results.

Self-check

  1. What does top-p = 0.9 mean?

  2. What should you do when reliable code output matters?

References

  1. Create a Message — sampling parameters — Anthropic

    Temperature and top-p limits; zero temperature is not a determinism guarantee.

  2. The Curious Case of Neural Text Degeneration — Holtzman et al. / arXiv

    Primary research introducing nucleus sampling.

Last verified 2026-08-19.