Positional information in transformers
Learn how transformers know word order, from sinusoids to rotary embeddings and long context limits
A taste of a lesson
Why does rotating the query and key give relative positions instead of absolute ones?
Take one two dimensional pair. Rotary embedding rotates the query at position m by angle m times theta, and the key at position n by n times theta. A dot product between two rotated vectors equals the dot product of the originals with one of them rotated by the difference of the angles. So the score depends on (m minus n) times theta, not on m and n separately. Shift both tokens ten places later and the score is unchanged. Quick check: if a query is at position 7 and a key at position 3, what angle difference enters the score for a pair with frequency theta?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Show why self attention without positional information ignores word order
- Compute sinusoidal encodings and explain learned absolute positions
- Explain how rotary embeddings make attention depend on relative distance
- Compare relative biases and distance penalties for longer inputs
- Judge context extension and long context claims with appropriate scepticism
Lesson plan
- 1 Why attention needs position Show that self attention treats its input as an unordered set. Start
- 2 Sinusoidal and learned absolute positions Compute fixed sinusoidal encodings and compare them with learned position vectors. Start
- 3 Relative position biases Add distance based biases directly to attention scores. Start
- 4 Rotary embeddings Work through how rotating queries and keys makes scores depend on relative distance. Start
- 5 Linear distance penalties Understand how subtracting a slope times distance biases heads towards nearby tokens. Start
- 6 Longer contexts and honest testing Explain context extension methods and how to test what a long context model really uses. Start
Try asking
About this tutor
For learners who know how attention works and have wondered how a transformer tells 'dog bites man' from 'man bites dog'. You start by proving that attention on its own ignores order, then work through the main ways models add positional information: sinusoidal encodings, learned absolute positions, relative position biases, rotary embeddings and linear distance penalties. For each you compute a small example and discuss what it implies for longer inputs. The final lessons cover why models often struggle beyond the length they were trained on, how context extension methods work in general terms, and how to test long context claims sensibly.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Architectures explained from the inside: convolutions, recurrence, attention and beyond
9 tutors 362 lessons taught Sample
I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...
See Nikolai's profile and tutorsMore like this
Other tutors on the same or nearby topics.