Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Attention without the maths

Understand how transformers let each word look at the others, explained with pictures and stories

Your first 3 replies from the tutor are free. Then you can continue the lesson for $5 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Attention without the maths AI tutor following Bastian Weber's plan
Student:

When people say the model 'pays attention', is it like me concentrating?

Tutor:

It is a borrowed word, so only loosely. In a transformer, attention is a step where every word looks at the other words and decides how much to take from each. In 'The cat sat on the mat because it was warm', the word 'it' takes a lot of information from 'mat' and a little from 'cat'. There is no focus or effort involved, just weighted blending, done for every word at once, many times over. In 'because it was hungry', which word should 'it' draw from most?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain in plain words what attention does inside a transformer
  • Describe why multiple attention heads and many layers are used
  • Explain why a model only attends to earlier words when generating text
  • Describe why long inputs are costly and what attention does not explain

Lesson plan

5 lessons. Pick one to start there.

  1. 1 Words need other words See why meaning depends on context, sometimes far away in a sentence. Start
  2. 2 The core idea: ask, match, gather Understand attention as each word gathering information from relevant words. Start
  3. 3 Many heads, many layers Learn why transformers run many attention processes and stack them. Start
  4. 4 Looking only backwards when writing Understand causal masking in text generation. Start
  5. 5 Costs and limits of attention Know why long contexts are expensive and what attention does not tell us. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For curious non technical learners who want to know what 'attention' means inside a language model without any equations. Using sentences, diagrams and everyday analogies, you see how each word gathers information from the other words that matter to it, how many attention heads look for different kinds of relationships at once, how layers stack these steps, and why a model only looks backwards when writing. You also learn why long inputs are expensive and what attention does not do, such as guarantee understanding. By the end you can explain the core idea of the transformer to a friend in a few clear sentences.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Bastian Weber

I explain how language models really work, from tokens to attention, without hand waving

9 tutors 4.5(24) 525 lessons taught Sample

I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...

See Bastian's profile and tutors