Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

The Chain Rule and Backpropagation

Compute gradients through a network by hand and see exactly what backpropagation does

By Leandro Ferraz Intermediate Math for AI 4.7(3) 52 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

The Chain Rule and Backpropagation AI tutor following Leandro Ferraz's plan
Student:

With p = w*x + b and L = (p - y)^2, where w=2, x=3, b=1, y=5, what is dL/dw?

Tutor:

Forward pass first. p = 2*3 + 1 = 7. Error p - y = 2, so L = 4. Now backward. The local derivative of L with respect to p is 2*(p - y) = 4. The local derivative of p with respect to w is x = 3. Chain rule: dL/dw = 4 * 3 = 12. So nudging w up by 0.01 should raise the loss by about 0.12. Check it: set w = 2.01, recompute L, and compare. Then try dL/db yourself. What do you get?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Apply the chain rule to nested functions confidently
  • Draw a computational graph and run forward and backward passes
  • Sum gradients correctly where paths branch
  • Explain vanishing and exploding gradients from the arithmetic
  • Verify gradients numerically with finite differences

Lesson plan

6 lessons. Pick one to start there.

  1. 1 The chain rule Differentiate nested functions by multiplying local rates. Start
  2. 2 Computational graphs Represent a calculation as a graph of simple operations. Start
  3. 3 One neuron by hand Run forward and backward passes for a linear model with squared loss. Start
  4. 4 Branches and layers Backpropagate through a two layer network. Start
  5. 5 Activations and gradient flow See how activations shape gradients through depth. Start
  6. 6 Autodiff and gradient checks Understand what frameworks automate and how to verify it. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

An intermediate tutor for learners who know basic derivatives and want to understand how neural networks compute gradients. You will master the chain rule, draw computational graphs, and run forward and backward passes by hand through a tiny model, then a two layer network. Lessons explain why gradients from branching paths add up, how activation functions affect gradient flow, why gradients can vanish or explode, and what automatic differentiation software is doing for you. You finish able to verify a gradient numerically and to reason about training problems from the maths.

Reviews

4.7

3 ratingsSample

  • Ayumi S.Sample

    The vanishing gradient explanation through the 0.25 sigmoid bound was simple and memorable.

  • Chukwudi E.Sample

    Clear and rigorous. The matrix form lesson needed a second pass, but the tutor slowed down when I asked.

  • Tobias R.Sample

    Doing backprop by hand with a table of values was the turning point for me. Branching gradients adding up now makes sense.

About the teacher

Leandro Ferraz

Calculus and optimisation for machine learning, done by hand

9 tutors 4.6(18) 319 lessons taught Sample

I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...

See Leandro's profile and tutors