Backpropagation in practice
Follow gradients through a real network and fix the training bugs that come from misusing them
A taste of a lesson
My loss is flat from the first step and the gradient norm for the first layers is exactly zero. What should I check?
Exactly zero, rather than tiny, usually means the graph is broken, not that gradients vanished. Check three things in order. First, is anything between the first layers and the loss detached or converted to a plain number, for example by moving it out of the framework's tensor type? Second, are those layers' parameters actually registered with the optimiser and set to require gradients? Third, does a ReLU layer receive only negative inputs, so every unit is dead? Print the gradient norm layer by layer to see where it becomes zero. Which layer is the first one with a nonzero gradient?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute gradients by hand for a small computation graph using the chain rule
- Explain how automatic differentiation records and replays a computation
- Diagnose vanishing and exploding gradients and choose sensible remedies
- Use gradient clipping, accumulation and numerical gradient checks correctly
- Find common bugs that stop gradients from reaching parameters
Lesson plan
- 1 The chain rule on a tiny graph Compute a full set of gradients by hand for a one weight model with squared loss. Start
- 2 Local gradients of common operations Learn how addition, multiplication, ReLU, sigmoid and matrix products pass gradients back. Start
- 3 Automatic differentiation in frameworks Connect hand computation to how frameworks record and replay the forward pass. Start
- 4 Vanishing and exploding gradients Explain why deep products of gradients misbehave and how architectures counter it. Start
- 5 Memory, accumulation and checkpointing Manage the memory cost of the backward pass and simulate larger batches. Start
- 6 Checking and debugging gradients Verify gradients numerically and locate bugs that block learning. Start
Try asking
About this tutor
For learners who know what a neural network is and have heard that backpropagation is 'just the chain rule', but want to see it work and use it confidently. You will compute gradients by hand for a tiny network, follow a computation graph forward and backward, and connect what you calculated to what automatic differentiation does inside any modern framework. The second half is practical: vanishing and exploding gradients, gradient clipping, gradient accumulation, memory use during the backward pass, checking custom gradients numerically, and the everyday bugs that stop a model from learning. Some comfort with derivatives helps; the tutor reviews what you need.
Reviews
4.5
4 ratingsSample
- Yusuf A.Sample
Strong course. The matrix gradient part went fast and I needed to ask for a second example, which the tutor gave without fuss.
- Arjun P.Sample
Lesson six's checklist found my bug: I was converting an intermediate tensor to a NumPy array for logging and reusing it. Gradients stopped right there.
- Hana K.Sample
I liked that it kept separating backprop from the optimiser. I had been mixing them up in interviews.
- Claire D.Sample
The napkin example made the chain rule feel obvious, then the course moved quickly into things I actually hit at work like accumulation and memory.
About the teacher
I teach how neural networks learn, one small worked example at a time
9 tutors 424 lessons taught Sample
I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...
See Mira's profile and tutorsMore like this
Other tutors on the same or nearby topics.