Feature Engineering for Tabular Data
Create features that help models learn, without leaking the answer
A taste of a lesson
My churn model has 'city' with 900 values. One hot encoding explodes. Can I replace city with its average churn rate?
Yes, that is target encoding, and it can work well, but done naively it leaks. If a row's own label contributes to its city's average, the model gets a hint of the answer, especially for small cities. Compute the encoding out of fold: for each fold, use averages from the other folds only. Also smooth rare cities towards the overall churn rate, for example by blending with a weight based on city size. What churn rate would a city with only three customers get without smoothing?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Design features from domain knowledge and the prediction moment
- Transform numeric features appropriately for linear and tree models
- Encode categorical features, including leak free target encoding
- Build date, cyclical and time window aggregate features
- Evaluate features by validation gains and remove weak ones
Lesson plan
- 1 Features from understanding Generate candidate features from how the process works. Start
- 2 Numeric transformations Reshape numeric features for the model family in use. Start
- 3 Encoding categories Choose an encoding for low and high cardinality categories. Start
- 4 Dates and cycles Extract useful time features. Start
- 5 Aggregates over time windows Summarise each entity's history without leaking the future. Start
- 6 Keeping what helps Evaluate and prune features. Start
Try asking
About this tutor
An intermediate tutor for learners who can train a basic model and want better inputs. You will transform numeric features, encode categories with one hot, frequency and out of fold target encoding, build date and cyclical features, compute per entity aggregates over time windows, add missing indicators and simple text features. Lessons explain how linear models and tree models need different features, and every technique comes with its leakage risk and how to avoid it. You practise on churn, demand and pricing style datasets and finish with a feature checklist.
Reviews
4.3
3 ratingsSample
- Annika S.Sample
Great time window aggregates lesson. The cyclical encoding trick was simple and useful.
- Karim J.Sample
Out of fold target encoding explained clearly at last. My validation scores dropped a bit but now match production.
- Sipho D.Sample
Practical and leak conscious. More on text features would be welcome.
About the teacher
Data cleaning, SQL, exploratory analysis and honest charts
9 tutors 427 lessons taught Sample
I teach the part of data science that takes most of the time: getting data into a shape you can trust, querying it, exploring it and showing it honestly. I came to data from operations work, where reports drove real decisions and a wrong join could cost a week. I teach by handing you small, deliberately messy tables and asking...
See Lin's profile and tutorsMore like this
Other tutors on the same or nearby topics.