Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Why some models are fast: efficiency tricks inside

Learn the techniques that make models faster and cheaper: from mixture of experts to KV caching

Your first 3 replies from the tutor are free. Then you can continue the lesson for $11 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Why some models are fast: efficiency tricks inside AI tutor following Benedict Asante's plan
Student:

If a mixture of experts model only uses a few experts per token, why does it still need so much memory?

Tutor:

Because the router can choose any expert for any token, so every expert must be ready in fast memory. Compute per token scales with the active experts, say two of sixteen, but memory scales with all of them. So an MoE model can generate with the compute of a much smaller dense model while needing the memory of a large one. Offloading experts to slower memory is possible but hurts speed. Exercise: with 16 experts of 2 billion parameters each and 2 active, what are total and active expert parameters?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Distinguish compute bound prefill from bandwidth bound decode
  • Estimate KV cache size and explain techniques that shrink it
  • Explain mixture of experts, quantisation and distillation and their trade offs
  • Describe speculative decoding, batching and prompt caching in serving

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Find the bottleneck Separate compute, memory capacity, bandwidth and latency limits. Start
  2. 2 The KV cache Understand and estimate the memory cost of caching attention state. Start
  3. 3 Mixture of experts See how sparse activation cuts compute per token. Start
  4. 4 Smaller numbers, smaller models Compare quantisation and distillation. Start
  5. 5 Speculative decoding Understand how a draft model speeds up generation. Start
  6. 6 Serving at scale Learn batching and caching strategies for many users. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For technical learners, engineers and analysts who want to understand how modern language models are made efficient to train and serve. You study mixture of experts architectures, the KV cache and why memory bandwidth limits generation, quantisation, distillation into smaller models, speculative decoding, batching strategies on servers, attention variants that cut memory, and prompt caching. For each, you learn the core idea, the trade off, and how it shows up in speed, cost and quality. Concepts are explained with simple numbers and diagrams in words, with no specific library or product details.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Benedict Asante

I explain the kinds of AI models, what they cost to run and how to run one yourself

9 tutors 4.4(16) 328 lessons taught Sample

I teach the practical side of modern models: reasoning models, multimodal models, open and closed weights, running a model on your own computer, and the money, energy and hardware behind every answer. I like starting with something you can see or measure, such as the memory a model needs or the number of tokens a task uses, then explaining the...

See Benedict's profile and tutors