Which Model When: Choosing an Algorithm
Pick a sensible model family for your data, constraints and goals, then test it fairly
A taste of a lesson
I have 3,000 rows of customer data with 25 features. Should I use a neural network to predict churn?
With 3,000 rows of tabular data, I would not start there. Begin with a baseline, such as predicting the overall churn rate, then a regularised logistic regression, which is interpretable and hard to beat on small data, and a gradient boosted tree model. Compare them with the same cross validation folds and a metric that fits your use, such as precision at the number of customers your team can contact. A neural network would likely need more data and tuning for little gain. Who will act on the predictions, and do they need reasons for each one?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Describe a problem by the factors that drive model choice
- Shortlist suitable model families with reasons
- Set a baseline and run a fair comparison of candidates
- Weigh accuracy against interpretability, latency, cost and maintenance
- Recognise when pretrained models or hosted language models make sense
Lesson plan
- 1 Describe the problem first Capture the facts that drive model choice. Start
- 2 Baselines and simple models Set the bar with baselines and interpretable models. Start
- 3 Tabular defaults Choose among linear models, forests and boosting. Start
- 4 Text, images and audio Pick approaches for unstructured data. Start
- 5 No labels or time order Handle unsupervised and forecasting problems. Start
- 6 A fair comparison Compare shortlisted models and decide. Start
Try asking
About this tutor
A decision coaching tutor for learners at any level who know some algorithms but freeze when choosing one. You will work through the factors that matter: data type and size, interpretability, latency and cost, the shape of the signal and how the model will be maintained. The tutor helps you shortlist candidates for your own problem, from regularised linear models and tree ensembles to pretrained neural networks or hosted language models, and design a fair comparison on the same validation scheme. Beginners get clear defaults; experienced learners discuss the trade offs in depth.
Reviews
4.3
3 ratingsSample
- Fatou S.Sample
Useful discussion of when a hosted language model is worth the cost for text classification. Balanced and practical.
- Tomas R.Sample
Talked me out of a neural network for a small tabular problem. Logistic regression won and my stakeholders love the explanations.
- Ravi K.Sample
Good defaults for a beginner, with deeper trade offs when I asked. The fair comparison checklist is now in my project template.
About the teacher
Probability for machine learning, plus forecasting and anomaly detection
9 tutors 361 lessons taught Sample
I teach probability the way machine learning uses it: random variables, likelihood, entropy and simulation. I also teach two applied areas where probability matters every day: time series forecasting and anomaly detection. My work background is in monitoring and forecasting for operational systems, where wrong alarms and missed incidents both have a cost. I teach through small simulations, coin and...
See Kenta's profile and tutorsMore like this
Other tutors on the same or nearby topics.