Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Clustering Without Guesswork

Group data with k-means, hierarchical clustering and DBSCAN, and check the groups are useful

By Katarzyna Wolska Beginner Machine learning 4.7(3) 33 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $4 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Clustering Without Guesswork AI tutor following Katarzyna Wolska's plan
Student:

I ran k-means on customers with k = 4 and got segments. How do I know four is right?

Tutor:

There is rarely a single right k, so gather several kinds of evidence. Plot within cluster error for k from 2 to 10 and look for a bend, though it is often vague. Compute silhouette scores for each k. Rerun with different random seeds and on resampled data: do the same segments keep appearing? Most important, profile each segment in plain words and ask whether marketing would treat them differently. If two segments would get the same treatment, merge them. How would you describe your four segments in one sentence each?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Run k-means by hand and explain each step
  • Choose the number of clusters using several kinds of evidence
  • Prepare features so distances reflect meaningful similarity
  • Apply hierarchical clustering and DBSCAN where k-means fails
  • Validate clusters by profiling, stability and usefulness

Lesson plan

6 lessons. Pick one to start there.

  1. 1 k-means by hand Run a few iterations of k-means on a small set of points. Start
  2. 2 Features and scaling Prepare features so distances mean something. Start
  3. 3 How many clusters Choose k with several kinds of evidence. Start
  4. 4 Hierarchical clustering Build and cut a dendrogram. Start
  5. 5 Density and soft clusters Use DBSCAN and mixture models when k-means fails. Start
  6. 6 Judging clusters Check clusters are stable and useful. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

A beginner tutor for anyone who wants to find natural groups in data, such as customer segments, document topics or sensor patterns. You will run k-means by hand on a few points, then learn how to choose the number of clusters, why scaling matters, and when k-means fails. Lessons introduce hierarchical clustering with dendrograms, density based clustering that finds noise points, and soft clustering with mixture models. Throughout, the tutor stresses that clusters are a tool rather than a discovered truth, and teaches you to judge them by stability and usefulness.

Reviews

4.7

3 ratingsSample

  • Amira L.Sample

    Our segmentation project finally has segments the sales team actually uses.

  • Dev R.Sample

    DBSCAN explanation was clear and helped with our sensor data. The elbow method honesty was refreshing.

  • Olga T.Sample

    Running k-means by hand on six points made the algorithm obvious. Profiling clusters in plain words was the most useful habit.

About the teacher

Katarzyna Wolska

Linear algebra for AI, with geometry first and notation second

9 tutors 4.6(21) 338 lessons taught Sample

I teach the linear algebra behind modern AI: vectors, matrices, similarity, eigenvectors and the methods built on them, such as PCA, clustering and recommender systems. I trained in applied mathematics and later worked on search and recommendation features, so I like to connect each idea to something a real system does. My lessons begin with pictures and small numbers you...

See Katarzyna's profile and tutors