Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Quantisation: smaller, faster models

Shrink models with lower precision numbers and measure exactly what quality you trade away

By Magnus Eriksen Intermediate Fine tuning and training MLOps and deployment 4.3(3) 49 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Quantisation: smaller, faster models AI tutor following Magnus Eriksen's plan
Student:

I quantised my model to 4 bit and it uses far less memory, but it is slower than the 16 bit version. How?

Tutor:

Memory savings follow directly from the arithmetic of bytes, but speed depends on how the computation runs. If your runtime stores weights in 4 bit and converts them back to 16 bit with a separate step before each multiplication, that extra work can outweigh the bandwidth saved. Fast setups use kernels that fuse dequantisation into the matrix multiplication. Also check batch size: at large batches you are compute bound, and weight only quantisation helps less. Benchmark both models at your real batch size and sequence length. Quick check: at batch size 1, is generation usually limited by arithmetic or by memory bandwidth?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Compare common number formats and their memory costs
  • Quantise and dequantise values by hand using scale and zero point
  • Explain how outliers and scale granularity affect quality
  • Choose between weight only, weight and activation, PTQ and QAT approaches
  • Measure quality loss and speed on your own task and hardware

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Number formats and memory Compare 32, 16, 8 and 4 bit formats and compute model sizes for each. Start
  2. 2 Scale and zero point by hand Map real values to integers and back, and measure the rounding error. Start
  3. 3 Granularity and outliers Understand why a few large values make low bit quantisation hard. Start
  4. 4 What to quantise and when Choose between weight only and weight and activation quantisation, and PTQ or QAT. Start
  5. 5 Measuring quality loss Evaluate a quantised model on the tasks that matter to you. Start
  6. 6 Speed, kernels and the cache Measure real speed and extend savings to the key value cache. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who want to run models with less memory and often more speed by storing numbers in lower precision. You will learn the common number formats, compute a quantised value and its error by hand with a scale and zero point, and see why outliers make low bit quantisation hard. The lessons cover per tensor, per channel and per group scales, weight only versus weight and activation quantisation, post training methods with calibration data versus quantisation aware training, and quantising the key value cache. Most importantly, you learn to measure quality loss on your own task, and why smaller does not always mean faster.

Reviews

4.3

3 ratingsSample

  • Ravi S.Sample

    Doing the scale and zero point example by hand made every later lesson easier. I now understand why group size appears in quantisation settings.

  • Kofi A.Sample

    Good balance of theory and practice. I would have liked more on QAT, but the tutor was clear it is less common for most deployments.

  • Elena V.Sample

    The point about testing maths and other languages separately was spot on. Our 4 bit model was fine in English and clearly worse in Polish.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors