Least Squares from First Principles
Derive the normal equations, see regression as projection and solve it the stable way
A taste of a lesson
Why are least squares residuals always orthogonal to the feature columns?
Because the fitted vector is the closest point to y within the space of all possible Xw, and the closest point in a flat space is reached by dropping a perpendicular. If the residual had any component along a feature column, you could move a little in that direction and get closer to y, so the fit would not be optimal. In symbols, X^T r = 0, which is exactly the normal equations. Check it on (1, 1), (2, 2), (3, 2) with y_hat = 0.5x + 2/3: do the residuals sum to zero?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Derive the normal equations using matrix calculus
- Interpret least squares as an orthogonal projection
- Explain how collinearity and conditioning affect solutions
- Justify solving with QR or SVD rather than an explicit inverse
- Connect least squares and ridge to maximum likelihood and MAP
Lesson plan
- 1 Setting up the problem Write least squares in matrix form with correct shapes. Start
- 2 The calculus route Derive the normal equations from the gradient. Start
- 3 The geometric route See fitted values as a projection onto the column space. Start
- 4 Conditioning and stable solving Understand collinearity and numerically stable solutions. Start
- 5 Ridge regression Derive ridge and see how it fixes ill conditioning. Start
- 6 The probabilistic route Link least squares to likelihood and priors. Start
Try asking
About this tutor
An advanced tutor for learners who use linear regression and want to understand its mathematics completely. You will derive the least squares solution with calculus, then again with geometry, seeing fitted values as the projection of the target onto the space spanned by the features. Lessons cover the simple regression formulas, the normal equations, why collinearity breaks the matrix inverse, why practitioners solve with QR or SVD instead of inverting, how ridge regression fixes ill conditioning, and how least squares arises as maximum likelihood under Gaussian noise. Expect matrix notation and careful derivations.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Calculus and optimisation for machine learning, done by hand
9 tutors 319 lessons taught Sample
I teach the calculus and optimisation that make models learn: derivatives, gradients, the chain rule, gradient descent and the loss functions it minimises. My background is in engineering and numerical computing, so I care about why things work and also about when they break, such as unstable learning rates or overflowing exponentials. I teach with pencil calculations first, then a...
See Leandro's profile and tutorsMore like this
Other tutors on the same or nearby topics.