Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

How language models are pretrained

Understand the data, objective, scaling and stability work behind large language model pretraining

By Mateo Rojas Advanced Deep learning NLP 4.7(3) 45 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $13 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

How language models are pretrained AI tutor following Mateo Rojas's plan
Student:

Roughly how much compute does it take to train a 7 billion parameter model on 2 trillion tokens?

Tutor:

Use the rule of thumb of about 6 operations per parameter per token: 6 times 7 times 10 to the 9 times 2 times 10 to the 12 gives about 8.4 times 10 to the 22 floating point operations. To turn that into time, divide by what your hardware sustains in practice, which is usually well below its peak figure because of communication and memory overheads. Treat the result as an order of magnitude estimate, not a budget. Your turn: if you doubled the tokens but halved the parameters, how would the compute change?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain what next token prediction teaches a model and what it does not
  • Describe a pretraining data pipeline including filtering, deduplication and mixing
  • Estimate training compute from parameter count and token count
  • Interpret scaling trends and the trade off between training and inference cost
  • Distinguish what pretraining provides from what post training adds

Lesson plan

6 lessons. Pick one to start there.

  1. 1 The next token objective Understand the pretraining loss and what kinds of knowledge it rewards. Start
  2. 2 Building the dataset Follow raw text through extraction, filtering and deduplication. Start
  3. 3 Tokenizers and data mixtures See how tokenizer choice and source weighting shape the model. Start
  4. 4 Compute and scaling trends Estimate training compute and interpret how loss scales with model and data size. Start
  5. 5 Keeping a long run stable Handle warmup, loss spikes, precision and checkpoints in large training runs. Start
  6. 6 From base model to assistant Separate what pretraining provides from what instruction and preference tuning add. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who know transformers and want to understand what actually happens when a large language model is pretrained. You will start from the next token prediction objective and what it does and does not teach, then follow the real work: collecting and filtering text, deduplication, choosing a tokenizer, mixing data sources, and checking for test set contamination. Then you study scaling: how loss improves with compute, parameters and data, what compute optimal means in general terms, and why many models are trained on more data than that. The final lessons cover training stability, mixed precision and what pretraining gives compared with later post training stages.

Reviews

4.7

3 ratingsSample

  • Mohammed J.Sample

    I liked that scaling laws were presented as empirical trends with limits, not magic. The stability lesson matched problems I had seen in smaller runs.

  • Shreya B.Sample

    Strong on data quality and contamination. I wish there was more on tokenizer training, but the tutor gave good pointers when I asked.

  • Anders L.Sample

    The 6ND estimate and the inference cost discussion explained why our team picked a smaller model trained longer. Very clear lessons on data pipelines too.

About the teacher

Mateo Rojas

Generative models and representation learning, explained with intuition first and maths second

9 tutors 4.6(20) 335 lessons taught Sample

I teach how models learn useful representations and how they generate new data: autoencoders, GANs, diffusion models, self supervised learning and language model pretraining. I came to this through research engineering work where we had to decide which kind of model was worth the compute, so I teach with trade offs in mind. Each topic starts with a picture or...

See Mateo's profile and tutors