Maximum Likelihood, Step by Step
Derive estimates by maximising likelihood and see why common losses are likelihoods in disguise
A taste of a lesson
Why is minimising mean squared error the same as maximum likelihood?
Assume each target equals the prediction plus Gaussian noise with a fixed variance. The probability density of one target is proportional to exp(-(y - y_hat)^2 / (2 sigma^2)). Multiply over examples and take the log: you get minus the sum of squared errors divided by 2 sigma^2, plus a constant. Maximising that is the same as minimising the sum of squared errors. So MSE is MLE under that noise assumption. Try it: what loss would you get if the noise followed a Laplace distribution, proportional to exp(-|error|)?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Distinguish probability from likelihood
- Write and maximise log likelihoods for simple models
- Derive MLEs for proportions, means and rates
- Show that MSE and cross entropy are negative log likelihoods
- Explain MLE overfitting and how MAP priors regularise
Lesson plan
- 1 Probability versus likelihood Understand likelihood as plausibility of parameters given data. Start
- 2 The coin example Write, log and maximise a likelihood by hand. Start
- 3 Gaussians and counts Derive MLEs for Gaussian and Poisson data. Start
- 4 Losses are likelihoods Connect MSE and cross entropy to likelihood. Start
- 5 When MLE overfits Recognise MLE failures with little data. Start
- 6 MAP and priors Add priors and see their link to regularisation. Start
Try asking
About this tutor
An intermediate tutor on the principle behind most model training. You will distinguish probability from likelihood, write likelihoods for coins, counts and Gaussian data, take logs and maximise by calculus, and derive familiar estimates such as the sample proportion and sample mean. Lessons then show that mean squared error and cross entropy are negative log likelihoods under specific assumptions, explain why maximum likelihood can overfit with little data, and introduce MAP estimation, where a prior acts like regularisation. Each derivation is checked numerically.
Reviews
4.7
3 ratingsSample
- Grace Y.Sample
Patient and methodical. The 0 heads in 3 flips example explained overfitting better than any textbook.
- Lena H.Sample
Seeing MSE fall out of the Gaussian likelihood was wonderful. Every loss function now has a story for me.
- Musa A.Sample
Clear step by step derivations. The MAP and regularisation link was the highlight.
About the teacher
Probability for machine learning, plus forecasting and anomaly detection
9 tutors 361 lessons taught Sample
I teach probability the way machine learning uses it: random variables, likelihood, entropy and simulation. I also teach two applied areas where probability matters every day: time series forecasting and anomaly detection. My work background is in monitoring and forecasting for operational systems, where wrong alarms and missed incidents both have a cost. I teach through small simulations, coin and...
See Kenta's profile and tutorsMore like this
Other tutors on the same or nearby topics.