Planning the cost of a training run
Estimate compute, time, memory and budget for a training or fine tuning run before you spend
A taste of a lesson
My manager wants a cost for fine tuning a 7 billion parameter model on 50 million tokens for 2 epochs. Where do I start?
Start with compute: 6 times parameters times tokens. Tokens seen are 50 million times 2 epochs, so 1e8. That gives 6 * 7e9 * 1e8 = 4.2e18 FLOPs. Next, pick an achieved throughput; as a hypothetical, 1.2e14 FLOPs per second. Divide: 4.2e18 / 1.2e14 = 35,000 seconds, under 10 GPU hours. Then check memory, because full fine tuning needs roughly 16 bytes per parameter, about 112 GB, so you may need several GPUs or adapters. Finally, add a buffer for evaluation and reruns. Your turn: what is the estimate if you only train for 1 epoch?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Estimate training compute with the 6 times parameters times tokens rule
- Convert compute into GPU hours using realistic utilisation
- Size memory for full and parameter efficient fine tuning
- Build a budget that includes failed runs, evaluations and data work
- Explain the trade offs of interruptible capacity and checkpointing
Lesson plan
- 1 The compute rule of thumb Estimate total training FLOPs from parameter count and training tokens. Start
- 2 From FLOPs to GPU hours Convert compute into time and cost using achieved, not peak, throughput. Start
- 3 Will it fit? Memory sizing Estimate memory for weights, gradients, optimiser state and activations. Start
- 4 Fine tuning versus pretraining Compare the scale of fine tuning runs with training from scratch. Start
- 5 The costs people forget Add realistic buffers for iteration, failure and supporting work. Start
- 6 Capacity choices and the full lifecycle Weigh interruptible capacity and compare training with ongoing inference cost. Start
Try asking
About this tutor
For engineers, researchers, managers and founders who need a defensible estimate of what a training or fine tuning run will cost. You will learn the standard rule of thumb for training compute, how to turn it into GPU hours using realistic utilisation rather than peak figures, and how to size memory so the run actually fits. The lessons compare pretraining with fine tuning, add the costs people forget, such as failed runs, evaluations and data work, and discuss cheaper interruptible capacity and its risks. Every example uses clearly hypothetical numbers, so you learn the method and plug in your own quotes.
Reviews
4.3
3 ratingsSample
- Patrick D.Sample
I presented a one page estimate with all assumptions listed and our finance lead approved it the same day. The buffer for failed runs turned out to be needed.
- Leila T.Sample
Very clear arithmetic. As a manager I skipped some of the memory detail, and the tutor adjusted without making me feel behind.
- Jakob M.Sample
The point about peak versus achieved throughput explained why our last run took twice as long as planned. Wanted more on multi node runs.
About the teacher
Making models fast, small and affordable: hardware, quantisation, serving and edge
9 tutors 355 lessons taught Sample
I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...
See Magnus's profile and tutorsMore like this
Other tutors on the same or nearby topics.