The attention formula, step by step
Work through queries, keys, values and softmax by hand until the transformer equation makes sense
A taste of a lesson
Why do we divide QK^T by the square root of d_k? It seems arbitrary.
It has a clear reason. If query and key entries are roughly independent with mean 0 and variance 1, their dot product sums d_k such terms, so its variance is about d_k and its typical size grows like sqrt(d_k). With d_k = 128, raw scores might differ by tens, and softmax then puts nearly all weight on one token, with gradients close to zero elsewhere. Dividing by sqrt(d_k) brings the variance back to about 1. Exercise: for d_k = 64, what number do we divide by?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute single head self attention by hand on a small example
- Explain the role of scaling, softmax and the causal mask
- Describe multi head attention and the cost of attention in sequence length
- Place attention inside a full transformer block with positions, residuals and normalisation
Lesson plan
- 1 Queries, keys and values Understand the three projections and what each represents. Start
- 2 Scores, scaling and softmax by hand Compute attention weights for a two token example. Start
- 3 Weighted values and the matrix form Finish the computation and express it compactly. Start
- 4 Masks and multiple heads Add causal masking and split attention into heads. Start
- 5 Cost, KV cache and positions Understand efficiency issues and why position must be added. Start
- 6 The full transformer block Assemble attention with normalisation, residuals and the feed forward layer. Start
Try asking
About this tutor
For technical learners comfortable with basic algebra who want to really understand self attention. You compute attention by hand on a tiny example: build query, key and value vectors, take dot products, scale, apply softmax, and form the weighted sum. Then you scale up to matrix form, add the causal mask, split into multiple heads, and see why the cost grows with the square of sequence length. You finish by placing attention inside a full transformer block with residual connections, normalisation and the feed forward layer, and by understanding why positional information has to be added. Short pseudocode is used, no specific library.
Reviews
4.5
2 ratingsSample
- Ilya S.Sample
Doing the two token example by hand made the formula obvious. The variance explanation for the scaling was the best I have seen.
- Neha G.Sample
Rigorous and well ordered. The KV cache and grouped query section was a bit quick for me, but the shape tracing exercise helped a lot.
About the teacher
I explain how language models really work, from tokens to attention, without hand waving
9 tutors 525 lessons taught Sample
I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...
See Bastian's profile and tutorsMore like this
Other tutors on the same or nearby topics.