Numerical Stability: Softmax and Log-Sum-Exp
Stop overflows, underflows and NaNs by computing ML formulas the stable way
A taste of a lesson
My softmax of logits [1000, 1001, 1002] returns [nan, nan, nan]. The maths seems fine. What's happening?
The maths is fine, the arithmetic is not. exp(1000) overflows to infinity in both float32 and float64, and infinity divided by infinity is NaN. Fix: subtract the maximum logit first. [1000, 1001, 1002] becomes [-2, -1, 0]. Softmax is unchanged by this shift, because the common factor exp(-1002) cancels in numerator and denominator, but now the exponents are tiny and safe. Compute it: exp(-2), exp(-1), exp(0), then normalise. What probabilities do you get?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain floating point range and precision limits in common formats
- Implement softmax and log-sum-exp without overflow
- Compute cross entropy stably from logits
- Avoid catastrophic cancellation in variance and similar computations
- Debug NaNs and infinities systematically, including in reduced precision
Lesson plan
- 1 How floating point fails Understand range, precision, overflow and underflow. Start
- 2 Stable softmax Derive and apply the max shift for softmax. Start
- 3 Log-sum-exp and log softmax Compute logs of sums of exponentials safely. Start
- 4 Sigmoids and small numbers Stabilise sigmoid, binary cross entropy and tiny value logs. Start
- 5 Cancellation and summation Avoid precision loss in subtraction and long sums. Start
- 6 Debugging and reduced precision Find numerical failures and train safely at lower precision. Start
Try asking
About this tutor
An advanced tutor for practitioners who implement or debug models and keep meeting infinities, NaNs and mysteriously poor training. You will learn how floating point numbers work and fail, then apply stable formulations of the most common ML computations: softmax with the max shift, the log-sum-exp trick, cross entropy computed from logits, stable sigmoid and log one plus x functions. Lessons also cover catastrophic cancellation, variance computations, safe epsilons and the special risks of reduced precision training. Each technique is derived and tested with small numbers you can check.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Calculus and optimisation for machine learning, done by hand
9 tutors 319 lessons taught Sample
I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...
See Leandro's profile and tutorsMore like this
Other tutors on the same or nearby topics.