Instruction tuning a base model
Turn a base model that only continues text into one that follows instructions in a reliable format
A taste of a lesson
After tuning, my model answers well but never stops. It keeps writing a fake user turn and answers itself. Why?
That is the classic sign that the model never learned when a turn ends. Two common causes. First, the end of turn token may be missing from the training targets, perhaps masked out along with the prompt or stripped by your formatting code. Second, inference may use a different chat template from training, so the stop marker the server waits for never appears. Render one training example exactly as the model sees it and check the final tokens, then compare with what your inference code sends. Quick check: in your rendered example, is the end of turn token inside the part that counts toward the loss?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain how base and instruction tuned models differ in behaviour
- Build a prompt and response dataset with a consistent chat template
- Apply loss masking, end of turn tokens and sequence packing correctly
- Plan data for diversity and quality rather than sheer volume
- Evaluate the tuned model for format, helpfulness and safety regressions
Lesson plan
- 1 Base models versus instruction models See how a base model continues text and why instructions alone do not steer it reliably. Start
- 2 Designing the dataset Write prompt and response pairs that teach the behaviour you actually want. Start
- 3 Chat templates and special tokens Serialise conversations consistently so training and inference match exactly. Start
- 4 Loss masking and packing Decide which tokens contribute to the loss and train efficiently on short examples. Start
- 5 Training settings and overfitting to style Choose epochs and learning rates that teach behaviour without copying quirks. Start
- 6 Evaluating helpfulness and safety Test the tuned model against the base model and check for weakened safety behaviour. Start
Try asking
About this tutor
For practitioners who want to understand what turns a pretrained base model into an assistant that follows instructions. You start by comparing how base and instruction tuned models respond to the same prompt, then build a supervised fine tuning dataset of prompt and response pairs. The lessons cover chat templates and special tokens, masking the loss on prompt tokens, packing sequences, how many examples you really need, and why diversity matters more than volume. You also learn what instruction tuning does not do: it teaches format and behaviour more than new facts, and it can weaken safety behaviour if you are careless. Each lesson ends with a short design exercise.
Reviews
4.0
3 ratingsSample
- Zainab O.Sample
Useful content, but I hoped for more on multi turn conversations. The safety regression lesson was the strongest part for me.
- Viktor H.Sample
Good on templates and masking. The point that tuning teaches format more than facts changed my plan; we added retrieval instead of more training data.
- Ines C.Sample
My model rambled endlessly and the culprit was exactly the missing end of turn token in the targets. The habit of rendering examples before training is now in my checklist.
About the teacher
Fine tuning with judgment: when to do it, how to do it well, and how to know it worked
9 tutors 428 lessons taught Sample
I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...
See Neha's profile and tutorsMore like this
Other tutors on the same or nearby topics.