Skip to content
SamplePreview build: teacher profiles, ratings, reviews and lesson counts are sample data.
All tutors

Instruction tuning a base model

Turn a base model that only continues text into one that follows instructions in a reliable format

By Neha Varadan Intermediate Fine tuning and training 4.0(3) 75 lessons taught Sample

Your first 3 replies from the tutor are free. Then you can continue the lesson for $9 from your credit.

A taste of a lesson

One question and the answer, as this tutor gives it.

Instruction tuning a base model AI tutor following Neha Varadan's plan
Student:

After tuning, my model answers well but never stops. It keeps writing a fake user turn and answers itself. Why?

Tutor:

That is the classic sign that the model never learned when a turn ends. Two common causes. First, the end of turn token may be missing from the training targets, perhaps masked out along with the prompt or stripped by your formatting code. Second, inference may use a different chat template from training, so the stop marker the server waits for never appears. Render one training example exactly as the model sees it and check the final tokens, then compare with what your inference code sends. Quick check: in your rendered example, is the end of turn token inside the part that counts toward the loss?

Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.

What you will be able to do

  • Explain how base and instruction tuned models differ in behaviour
  • Build a prompt and response dataset with a consistent chat template
  • Apply loss masking, end of turn tokens and sequence packing correctly
  • Plan data for diversity and quality rather than sheer volume
  • Evaluate the tuned model for format, helpfulness and safety regressions

Lesson plan

6 lessons. Pick one to start there.

  1. 1 Base models versus instruction models See how a base model continues text and why instructions alone do not steer it reliably. Start
  2. 2 Designing the dataset Write prompt and response pairs that teach the behaviour you actually want. Start
  3. 3 Chat templates and special tokens Serialise conversations consistently so training and inference match exactly. Start
  4. 4 Loss masking and packing Decide which tokens contribute to the loss and train efficiently on short examples. Start
  5. 5 Training settings and overfitting to style Choose epochs and learning rates that teach behaviour without copying quirks. Start
  6. 6 Evaluating helpfulness and safety Test the tuned model against the base model and check for weakened safety behaviour. Start

Try asking

Tap a question to start a lesson with it.

About this tutor

For practitioners who want to understand what turns a pretrained base model into an assistant that follows instructions. You start by comparing how base and instruction tuned models respond to the same prompt, then build a supervised fine tuning dataset of prompt and response pairs. The lessons cover chat templates and special tokens, masking the loss on prompt tokens, packing sequences, how many examples you really need, and why diversity matters more than volume. You also learn what instruction tuning does not do: it teaches format and behaviour more than new facts, and it can weaken safety behaviour if you are careless. Each lesson ends with a short design exercise.

Reviews

4.0

3 ratingsSample

  • Zainab O.Sample

    Useful content, but I hoped for more on multi turn conversations. The safety regression lesson was the strongest part for me.

  • Viktor H.Sample

    Good on templates and masking. The point that tuning teaches format more than facts changed my plan; we added retrieval instead of more training data.

  • Ines C.Sample

    My model rambled endlessly and the culprit was exactly the missing end of turn token in the targets. The habit of rendering examples before training is now in my checklist.

About the teacher

Neha Varadan

Fine tuning with judgment: when to do it, how to do it well, and how to know it worked

9 tutors 4.6(21) 428 lessons taught Sample

I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...

See Neha's profile and tutors