Random Forests and Bagging
Understand why averaging many trees works and how to tune a forest sensibly
A taste of a lesson
If one deep tree overfits, why doesn't a forest of 500 deep trees overfit 500 times more?
Because each tree overfits in a different direction. Each one sees a different bootstrap sample and different feature choices at each split, so its noisy mistakes are partly random. When you average 500 of them, those random errors tend to cancel, while the real pattern they share remains. That is variance reduction. It only works if trees are not too similar, which is why random feature selection matters. Quick check: what would happen to this benefit if every tree used the same rows and all features?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain how bootstrap sampling and averaging reduce variance
- Use out of bag error as a validation estimate and know its limits
- Tune the few random forest settings that matter most
- Recognise extrapolation and calibration problems in forests
- Interpret feature importance with permutation methods
Lesson plan
- 1 The wisdom of noisy voters See why averaging many high variance models produces a stable prediction. Start
- 2 Bootstrap samples and bagging Understand how bagging creates diverse trees from one dataset. Start
- 3 Random feature selection Learn how restricting features per split decorrelates trees. Start
- 4 Tuning without fuss Set number of trees, leaf size and depth sensibly. Start
- 5 Limits of forests Recognise where forests fail: extrapolation, calibration and size. Start
- 6 Importance done properly Interpret feature importance without common biases. Start
Try asking
About this tutor
An intermediate tutor for learners who understand single decision trees and want to know why forests are so dependable. You will see how bootstrap sampling and random feature selection turn unstable trees into a stable ensemble, use out of bag error as a free validation check, and learn which settings actually matter. Lessons also cover the limits: forests cannot extrapolate, their probabilities may need calibration, and their default feature importances can mislead. Exercises use small simulations and realistic tabular problems, and you finish able to defend a random forest choice to a sceptical colleague.
Reviews
4.3
3 ratingsSample
- Chen W.Sample
Practical and calm. The 'crowd of noisy voters' framing stuck with me.
- Viktor H.Sample
Clear explanation of OOB error. Would have liked a comparison with boosting, but I see there is a separate tutor for that.
- Ines M.Sample
The extrapolation lesson explained exactly why my forest flatlined on a growing sales series. Permutation importance is now my default.
About the teacher
Classical machine learning models, worked through on paper before any code
9 tutors 379 lessons taught Sample
I teach the classical machine learning models: regression, trees, ensembles, nearest neighbours, support vector machines and naive Bayes. My background is in applied analytics, where I spent a lot of time explaining to colleagues why a model made a particular prediction, and that shaped how I teach. I like to start with a tiny dataset you can hold in your...
See Kavya's profile and tutorsMore like this
Other tutors on the same or nearby topics.