Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Scaling laws and where AI progress comes from

Understand how compute, data and model size drive progress, and the debates about what comes next

By Benedict Asante Advanced How language models work 4.5(2) 33 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $10 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Scaling laws and where AI progress comes from AI tutor following Benedict Asante's plan
Student:

If a model has 10 billion parameters, how many tokens should it train on to be compute optimal?

Tutor:

Using the rough compute optimal ratio of about 20 tokens per parameter, that is around 200 billion tokens. Training compute would then be roughly 6 times 10 billion times 200 billion, which is 1.2 times 10 to the 22 FLOPs. Two caveats: the ratio comes from a specific study's setup, and in practice developers often train smaller models on many more tokens because they are cheaper to run for millions of users. Exercise: compute the optimal tokens and FLOPs for a 70 billion parameter model.

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain empirical scaling laws relating loss to parameters, data and compute
  • Apply the compute estimate and compute optimal ratio in simple calculations
  • Assess the debate over emergent abilities and measurement
  • Explain test time compute as an additional scaling axis
  • Critically evaluate forecasts about future AI progress

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Power laws and loss Understand the empirical relationship between scale and loss. Start
  2. 2 Compute and the optimal balance Calculate training compute and the compute optimal split. Start
  3. 3 Emergence or measurement? Evaluate claims that abilities appear suddenly with scale. Start
  4. 4 Efficiency, data and synthetic data Understand progress beyond raw scale. Start
  5. 5 Test time compute See how inference time computation adds a new scaling axis. Start
  6. 6 Reading forecasts Judge predictions about where AI is heading. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For technically minded learners, analysts and policy readers who want to understand why AI capabilities have grown and how to think about future progress. You study the empirical scaling laws relating loss to compute, parameters and data, the compute optimal training result, the role of algorithmic efficiency and data quality, emergent abilities and the dispute over whether they are real or measurement artefacts, and the newer axis of test time compute. You also examine the main constraints and debates: data limits, synthetic data, energy and cost, and whether scaling continues to deliver. You learn to read forecasts critically.

Reviews

4.5

2 ratingsSample

  • Hiroshi T.Sample

    The worked compute calculations made the papers readable. The emergence debate was presented fairly with both sides.

  • Natalia Z.Sample

    Rigorous and well paced. The forecast reading lesson was useful for my policy work. Data limits section could go deeper.

About the teacher

Benedict Asante

I explain the kinds of AI models, what they cost to run and how to run one yourself

9 tutors 4.4(16) 328 lessons taught Sample

I teach the practical side of modern models: reasoning models, multimodal models, open and closed weights, running a model on your own computer, and the money, energy and hardware behind every answer. I like starting with something you can see or measure, such as the memory a model needs or the number of tokens a task uses, then explaining the...

See Benedict's profile and tutors