Attention, Demystified: From Dot Products to Explaining Transformers
Week 2, Part 2: The Full Model and the Explanation

Training and Generation in Practice

Explain next-token prediction with cross-entropy loss, and describe how a trained model generates text one token at a time, including the role of temperature in sampling.