Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Cross Validation Without Fooling Yourself

Use k fold, grouped, time series and nested cross validation correctly

By Lukas Brenner Intermediate Machine learning 4.5(2) 34 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $6 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Cross Validation Without Fooling Yourself AI tutor following Lukas Brenner's plan
Student:

I selected the top 20 features using the full dataset, then ran 5 fold cross validation. Score is great. Problem?

Tutor:

Yes, unfortunately. Your feature selection saw the labels of every row, including the rows that later act as validation folds. The chosen features are partly tuned to those rows, so the cross validation score is optimistic, sometimes dramatically, especially with many features and few rows. Put selection inside a pipeline so it is redone on each training fold. Then compare the new score with the old one. Before you run it, what do you expect: a small drop, a large drop, or no change?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Run k fold cross validation and report mean and spread
  • Choose stratified, grouped or time series schemes to match the data
  • Keep every preprocessing step inside the folds with a pipeline
  • Explain why tuned cross validation scores are optimistic
  • Design a nested cross validation for honest estimates

Lesson plan

6 lessons. Pick one to start there.

  1. 1 k fold, step by step Walk through k fold cross validation and report results properly. Start
  2. 2 Matching the scheme to the data Pick stratified, grouped or time series folds for the situation. Start
  3. 3 Pipelines inside folds Prevent leakage by fitting preprocessing within each training fold. Start
  4. 4 Reading variation Interpret fold to fold spread and decide when differences matter. Start
  5. 5 Tuning and optimism Understand why the best tuned score overstates real performance. Start
  6. 6 A validation plan for a project Write a validation plan a reviewer could check. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

An intermediate tutor for learners who know about train and test splits and now want reliable model comparisons. You will run k fold cross validation, read the spread of scores as well as the mean, and choose the right variant for your data: stratified, grouped or time series. Lessons explain why preprocessing must sit inside each fold, why scores after heavy tuning are optimistic, and how nested cross validation gives an honest estimate. You finish able to design a validation plan for a real project and explain it to a reviewer.

Reviews

4.5

2 ratingsSample

  • Jonas E.Sample

    Nested cross validation finally makes sense. A diagram would have helped, but the step by step walkthrough worked.

  • Adaeze O.Sample

    Feature selection outside the folds was exactly my mistake. My score dropped a lot once I fixed it, and now I trust it.

About the teacher

Lukas Brenner

Model evaluation you can trust: splits, metrics, leakage and tuning

9 tutors 4.5(22) 439 lessons taught Sample

Most of the machine learning failures I have seen were not about the algorithm. They came from a test set that was used too often, a feature that leaked the answer, or a metric that hid the real cost of mistakes. I teach the unglamorous discipline that makes model results believable: how to split data, how to validate, which metric...

See Lukas's profile and tutors