Batching and caching for model inference
Serve more requests on the same hardware by batching smartly and caching what can be reused
A taste of a lesson
Our GPU utilisation looks low, yet we cannot add more concurrent users without running out of memory. Why?
That usually means the key value cache, not compute, is your limit. Each active request holds cache memory for every token in its context, and long prompts or long answers add up fast. Estimate it: 2 times layers times key value heads times head dimension times bytes per value gives memory per token; multiply by typical context length and concurrent users. Remedies include paged cache allocation to cut fragmentation, prefix caching for shared system prompts, a quantised cache, or capping context length. Quick exercise: with 128 KB per token and 3,000 token contexts, how many concurrent requests fit in 24 GB of free memory?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Distinguish latency, time to first token and throughput
- Compare static, dynamic and continuous batching for different workloads
- Estimate key value cache memory and its effect on concurrency
- Apply prefix and response caching while managing staleness and privacy
- Load test and read tail latency percentiles correctly
Lesson plan
- 1 Latency versus throughput Define the metrics that matter for inference and how batching trades them. Start
- 2 Static, dynamic and continuous batching Choose a batching strategy that fits your request pattern and model type. Start
- 3 KV cache memory Estimate cache size per request and see how it limits concurrency. Start
- 4 Prefix caching Reuse computation for shared prompt prefixes across requests. Start
- 5 Response and semantic caching Cache full answers where safe and avoid stale or leaked responses. Start
- 6 Queues, tail latency and load testing Predict and measure behaviour under realistic load. Start
Try asking
About this tutor
For engineers running model inference who want more throughput without letting latency get out of hand. You will separate latency from throughput, compare static, dynamic and continuous batching, and see why continuous batching matters so much for language models whose responses vary in length. You estimate key value cache memory per request, learn how paged cache allocation and prefix caching work in principle, and weigh response caching and semantic caching against staleness and privacy risks. The course finishes with queueing behaviour, tail latency and how to load test with realistic traffic. Examples stay tool neutral.
Reviews
4.7
3 ratingsSample
- Maryam K.Sample
The privacy point about response caches was important for us. We had been caching by question text only, across users.
- Bruno T.Sample
Load testing with real output length distributions changed our capacity plan completely. Our earlier tests used fixed short answers.
- Sven L.Sample
Estimating cache memory per request explained our concurrency ceiling in ten minutes. We moved the system prompt to the front and prefix caching started working.
About the teacher
Making models fast, small and affordable: hardware, quantisation, serving and edge
9 tutors 355 lessons taught Sample
I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...
See Magnus's profile and tutorsMore like this
Other tutors on the same or nearby topics.