Gradient Descent by Hand
Take gradient descent steps with a pencil and feel how learning rates make or break training
A taste of a lesson
With f(w) = (w - 3)^2, start w = 0 and learning rate 1.0, my w goes 0, 6, 0, 6... What's wrong?
Nothing is wrong with your arithmetic; the learning rate is exactly too big. At w = 0 the gradient is -6, and stepping 1.0 times that moves you to 6, the same distance past the minimum on the other side. At 6 the gradient is +6, so you jump back to 0. Each step overshoots by exactly the gap. Try learning rate 0.4 and fill in three steps. Before you calculate, predict: will you approach 3 smoothly, or bounce around it while getting closer?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Perform gradient descent updates by hand
- Predict the effect of small, good and too large learning rates
- Extend descent to two parameters and fit a line to data
- Explain how feature scaling changes convergence
- Compare full batch, stochastic and mini batch updates
Lesson plan
- 1 Walking downhill Perform gradient descent updates on a one parameter function. Start
- 2 Choosing the learning rate See slow, fast, oscillating and diverging behaviour. Start
- 3 Two parameters Update two parameters using partial derivatives. Start
- 4 Fitting a line Use gradient descent to fit a line to a few data points. Start
- 5 Scaling features Understand why scaling speeds up convergence. Start
- 6 Batches and stopping Compare update styles and decide when to stop. Start
Try asking
About this tutor
A beginner tutor that makes gradient descent tangible. You will minimise simple functions by hand, step by step, choose learning rates that work, and watch what happens when they are too small, too large or wildly too large. Lessons then extend the idea to two parameters, to fitting a line to data, and to the difference between full batch, stochastic and mini batch updates. You will also see why feature scaling changes how quickly descent converges. Each lesson includes a short table of steps to fill in with a calculator.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Calculus and optimisation for machine learning, done by hand
9 tutors 319 lessons taught Sample
I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...
See Leandro's profile and tutorsMore like this
Other tutors on the same or nearby topics.