Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

The transformer, block by block

Trace a token through every part of a transformer and count where the parameters live

By Nikolai Sorin Advanced Deep learning How language models work 4.7(3) 60 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

The transformer, block by block AI tutor following Nikolai Sorin's plan
Student:

How many parameters does one block have if the width is 1024 and the FFN expands 4x?

Tutor:

Split it into the two sublayers. Attention has four projection matrices, Q, K, V and output, each 1024 by 1024, so 4 times 1,048,576, about 4.2 million. The feed forward network goes 1024 to 4096 and back, two matrices of 1024 by 4096, so 8 times 1,048,576, about 8.4 million. Together that is about 12.6 million, which is the 12d squared rule, ignoring small biases and norm weights. Notice the feed forward part is twice the attention. Now you try: with 24 such blocks and a 50,000 token vocabulary with tied embeddings, what is the rough total?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Trace a token through embedding, residual stream, attention, feed forward and output head
  • Estimate a transformer's parameter count from its width, depth and vocabulary
  • Explain pre norm residual blocks and why the residual stream matters
  • Compare encoder only, decoder only and encoder decoder designs and their uses
  • Read a model configuration file and explain each architectural setting

Lesson plan

6 lessons. Pick one to start there.

  1. 1 From token ids to vectors Turn token ids into embedding vectors and understand the size of the embedding table. Start
  2. 2 The residual stream and pre norm blocks See how each block reads from and writes back to a shared stream of vectors. Start
  3. 3 Attention inside the block Place multi head attention in the block and track its shapes and parameters. Start
  4. 4 The feed forward network Understand the per position MLP that holds most of each block's parameters. Start
  5. 5 Stacking blocks and the output head Go from the final hidden state to next token probabilities. Start
  6. 6 Encoder only, decoder only, encoder decoder Match the three transformer families to the tasks they suit. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For learners who already understand attention and want to see the whole transformer as one working machine. You follow a single token from its embedding through the residual stream, normalisation, multi head attention and the feed forward network, then stack the blocks and reach the output head that turns vectors back into word probabilities. Along the way you track shapes, estimate parameter counts for a realistic model, and compare encoder only, decoder only and encoder decoder designs. The aim is that you can read any transformer architecture diagram or config file and say what each part does, how big it is and why it is there.

Reviews

4.7

3 ratingsSample

  • Viktor H.Sample

    The 12d squared rule plus the embedding table let me sanity check a config file at work in about a minute. Lesson two on the residual stream changed how I picture the whole model.

  • Tariq S.Sample

    Finally understand why people say the MLP layers store a lot of the knowledge. Counting parameters by hand made it obvious where the bulk is.

  • Ines M.Sample

    Dense but well ordered. I would have liked one more example of a gated feed forward layer, though the tutor explained the parameter trade off clearly when I asked.

About the teacher

Nikolai Sorin

Architectures explained from the inside: convolutions, recurrence, attention and beyond

9 tutors 4.6(22) 362 lessons taught Sample

I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...

See Nikolai's profile and tutors