Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Catastrophic forgetting and how to limit it

Measure what a model loses when you fine tune it, and use replay, regularisation and merging to limit it

Your first 3 replies from the tutor are free. Then you can continue the lesson for $12 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Catastrophic forgetting and how to limit it AI tutor following Neha Varadan's plan
Student:

After full fine tuning on legal documents, our model's general writing got noticeably worse. What should I try first?

Tutor:

First, measure it so you can compare fixes: run a fixed general writing sample and your legal task at each saved checkpoint. Often the early checkpoints keep most of the legal gain with far less damage, which tells you to train fewer steps or at a lower learning rate. Next, try replay: mix a share of general writing data into each batch and tune that share on validation. If you need more, compare against a LoRA run or interpolate between base and tuned weights. Change one thing at a time. Quick check: what would it suggest if even your first checkpoint already shows the damage?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain why fine tuning overwrites earlier abilities
  • Measure forgetting with before and after suites across checkpoints
  • Apply replay, lower learning rates and parameter efficient methods deliberately
  • Compare weight, Fisher weighted and output space regularisation
  • Use weight interpolation to trade target gains against retained ability

Lesson plan

6 lessons. Pick one to start there.

  1. 1 What forgetting is and why it happens Connect gradient updates on narrow data to the loss of earlier abilities. Start
  2. 2 Measuring forgetting properly Build before and after suites and track them across checkpoints. Start
  3. 3 Training choices: learning rate, steps and data Reduce forgetting through gentler optimisation and data replay. Start
  4. 4 Regularising towards the original model Compare penalties on weights and on outputs that keep the model near its start. Start
  5. 5 Adapters, merging and interpolation Use architecture and weight space tools to contain or trade off forgetting. Start
  6. 6 The continual learning view Place these methods in the stability and plasticity trade off and its open problems. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For experienced practitioners whose fine tuned models got better at the target task and worse at almost everything else. You will learn why gradient updates on new data overwrite earlier abilities, how to measure forgetting with before and after suites, and which training choices make it worse. Then you work through mitigations with their trade offs: lower learning rates and shorter training, replaying general data, parameter efficient methods, penalties that keep weights or outputs close to the original model, and merging tuned and base weights. The course frames all of this as the stability and plasticity trade off from continual learning research, with honest notes on what remains unsolved.

Reviews

Students can review a tutor after a paid lesson. Nobody has yet.

About the teacher

Neha Varadan

Fine tuning with judgment: when to do it, how to do it well, and how to know it worked

9 tutors 4.6(21) 428 lessons taught Sample

I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...

See Neha's profile and tutors