The Chain Rule and Backpropagation
Compute gradients through a network by hand and see exactly what backpropagation does
A taste of a lesson
With p = w*x + b and L = (p - y)^2, where w=2, x=3, b=1, y=5, what is dL/dw?
Forward pass first. p = 2*3 + 1 = 7. Error p - y = 2, so L = 4. Now backward. The local derivative of L with respect to p is 2*(p - y) = 4. The local derivative of p with respect to w is x = 3. Chain rule: dL/dw = 4 * 3 = 12. So nudging w up by 0.01 should raise the loss by about 0.12. Check it: set w = 2.01, recompute L, and compare. Then try dL/db yourself. What do you get?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Apply the chain rule to nested functions confidently
- Draw a computational graph and run forward and backward passes
- Sum gradients correctly where paths branch
- Explain vanishing and exploding gradients from the arithmetic
- Verify gradients numerically with finite differences
Lesson plan
- 1 The chain rule Differentiate nested functions by multiplying local rates. Start
- 2 Computational graphs Represent a calculation as a graph of simple operations. Start
- 3 One neuron by hand Run forward and backward passes for a linear model with squared loss. Start
- 4 Branches and layers Backpropagate through a two layer network. Start
- 5 Activations and gradient flow See how activations shape gradients through depth. Start
- 6 Autodiff and gradient checks Understand what frameworks automate and how to verify it. Start
Try asking
About this tutor
An intermediate tutor for learners who know basic derivatives and want to understand how neural networks compute gradients. You will master the chain rule, draw computational graphs, and run forward and backward passes by hand through a tiny model, then a two layer network. Lessons explain why gradients from branching paths add up, how activation functions affect gradient flow, why gradients can vanish or explode, and what automatic differentiation software is doing for you. You finish able to verify a gradient numerically and to reason about training problems from the maths.
Reviews
4.7
3 ratingsSample
- Ayumi S.Sample
The vanishing gradient explanation through the 0.25 sigmoid bound was simple and memorable.
- Chukwudi E.Sample
Clear and rigorous. The matrix form lesson needed a second pass, but the tutor slowed down when I asked.
- Tobias R.Sample
Doing backprop by hand with a table of values was the turning point for me. Branching gradients adding up now makes sense.
About the teacher
Calculus and optimisation for machine learning, done by hand
9 tutors 319 lessons taught Sample
I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...
See Leandro's profile and tutorsMore like this
Other tutors on the same or nearby topics.