Loss functions: what your model is minimising
Understand MSE, cross entropy and friends well enough to choose, read and debug them
A taste of a lesson
My 10-class image model starts training with a loss of about 9. Is that normal?
Probably not. If the model started out guessing evenly across 10 classes, cross entropy would be about ln 10, which is 2.3. A starting loss near 9 means the model is very confident and very wrong from the first step. Common causes: labels that do not match the inputs, a softmax applied twice or not at all, unnormalised inputs producing huge logits, or an output layer initialised with large weights. Print a few predicted logits and check their size. Quick exercise: what starting loss would you expect for a 2-class problem?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute MSE, MAE and cross entropy for small examples by hand
- Choose a loss that fits the task and the kind of errors that matter
- Use the expected starting loss as a quick sanity check on a new model
- Explain why training loss and reported metrics can disagree
Lesson plan
- 1 What a loss function is for Understand the loss as the single number training tries to reduce. Start
- 2 Losses for predicting numbers Compare MSE, MAE and Huber and choose between them for a regression task. Start
- 3 Cross entropy for classification Compute binary and categorical cross entropy and see how they punish confident errors. Start
- 4 Imbalanced classes and weighted losses Adjust the loss when some classes are much rarer than others. Start
- 5 Loss versus metric, and reading loss curves Interpret loss values during training and relate them to the metrics you report. Start
Try asking
About this tutor
For beginners who can train a simple model but treat the loss as a black box number. You will compute mean squared error and cross entropy by hand, see how each one punishes different kinds of mistakes, and learn what a sensible loss value looks like at the start of training. The lessons cover regression losses and their sensitivity to outliers, binary and multi class cross entropy, weighted losses for imbalanced classes, and why the loss you train on is often different from the metric you report. Each lesson includes a small calculation and a realistic debugging question so the ideas stick.
Reviews
4.0
3 ratingsSample
- Ngozi E.Sample
The ln K sanity check alone was worth it. I now look at the first logged loss before anything else and it caught a label shuffle bug last week.
- Erik S.Sample
Solid basics but I already knew most of the regression part. The cross entropy and imbalance lessons were the useful bits for me.
- Pablo G.Sample
Clear and practical. The section on loss versus metric answered a question I had been confused about for months.
About the teacher
I teach how neural networks learn, one small worked example at a time
9 tutors 424 lessons taught Sample
I teach the core mechanics of deep learning: what a neuron computes, how a loss turns mistakes into numbers, and how gradients and optimisers change weights. My background is in building and training models for applied research teams, which mostly meant staring at loss curves that refused to go down. That shaped how I teach. I start every idea with...
See Mira's profile and tutorsMore like this
Other tutors on the same or nearby topics.