Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Planning the cost of a training run

Estimate compute, time, memory and budget for a training or fine tuning run before you spend

By Magnus Eriksen All levels Fine tuning and training MLOps and deployment 4.3(3) 65 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $7 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Planning the cost of a training run AI tutor following Magnus Eriksen's plan
Student:

My manager wants a cost for fine tuning a 7 billion parameter model on 50 million tokens for 2 epochs. Where do I start?

Tutor:

Start with compute: 6 times parameters times tokens. Tokens seen are 50 million times 2 epochs, so 1e8. That gives 6 * 7e9 * 1e8 = 4.2e18 FLOPs. Next, pick an achieved throughput; as a hypothetical, 1.2e14 FLOPs per second. Divide: 4.2e18 / 1.2e14 = 35,000 seconds, under 10 GPU hours. Then check memory, because full fine tuning needs roughly 16 bytes per parameter, about 112 GB, so you may need several GPUs or adapters. Finally, add a buffer for evaluation and reruns. Your turn: what is the estimate if you only train for 1 epoch?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Estimate training compute with the 6 times parameters times tokens rule
  • Convert compute into GPU hours using realistic utilisation
  • Size memory for full and parameter efficient fine tuning
  • Build a budget that includes failed runs, evaluations and data work
  • Explain the trade offs of interruptible capacity and checkpointing

Lesson plan

6 lessons. Pick one to start there.

  1. 1 The compute rule of thumb Estimate total training FLOPs from parameter count and training tokens. Start
  2. 2 From FLOPs to GPU hours Convert compute into time and cost using achieved, not peak, throughput. Start
  3. 3 Will it fit? Memory sizing Estimate memory for weights, gradients, optimiser state and activations. Start
  4. 4 Fine tuning versus pretraining Compare the scale of fine tuning runs with training from scratch. Start
  5. 5 The costs people forget Add realistic buffers for iteration, failure and supporting work. Start
  6. 6 Capacity choices and the full lifecycle Weigh interruptible capacity and compare training with ongoing inference cost. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For engineers, researchers, managers and founders who need a defensible estimate of what a training or fine tuning run will cost. You will learn the standard rule of thumb for training compute, how to turn it into GPU hours using realistic utilisation rather than peak figures, and how to size memory so the run actually fits. The lessons compare pretraining with fine tuning, add the costs people forget, such as failed runs, evaluations and data work, and discuss cheaper interruptible capacity and its risks. Every example uses clearly hypothetical numbers, so you learn the method and plug in your own quotes.

Reviews

4.3

3 ratingsSample

  • Patrick D.Sample

    I presented a one page estimate with all assumptions listed and our finance lead approved it the same day. The buffer for failed runs turned out to be needed.

  • Leila T.Sample

    Very clear arithmetic. As a manager I skipped some of the memory detail, and the tutor adjusted without making me feel behind.

  • Jakob M.Sample

    The point about peak versus achieved throughput explained why our last run took twice as long as planned. Wanted more on multi node runs.

About the teacher

Magnus Eriksen

Making models fast, small and affordable: hardware, quantisation, serving and edge

9 tutors 4.6(18) 355 lessons taught Sample

I teach the engineering side of running models: what GPUs actually do, how memory limits shape every decision, how to quantise or distil a model, and how to serve it efficiently on a server or a small device. I have spent my working life close to hardware, first on embedded systems and later on inference infrastructure, so I tend to...

See Magnus's profile and tutors