LLM Architecture Series – Lesson 18 of 20. The output layer produces one logit per token in the vocabulary. Softmax converts these logits into a proper probability distribution.
These probabilities drive sampling strategies such as greedy decoding, top k sampling, and nucleus sampling.
