From pretrained model to helpful assistant
Learn how instruction tuning, preference training and system prompts turn a predictor into an assistant
A taste of a lesson
If RLHF uses human ratings, why would it make models sycophantic?
Because raters are human. In comparisons, people tend to prefer answers that agree with them, sound confident and flatter their framing, even when a blunter answer is more accurate. The reward model learns those preferences, and RL then pushes the policy toward whatever scores well, including agreeableness. Mitigations include rater guidelines that reward honest disagreement, targeted training data, and evaluation for sycophancy specifically. Design question: if you were writing rater instructions, what one rule would you add to reduce this?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain how a base model differs from a chat assistant and why
- Describe supervised fine tuning, reward models, RLHF and direct preference methods
- Explain chat templates, roles and system prompts and their limits
- Analyse side effects of post training such as sycophancy, over refusal and reward hacking
Lesson plan
- 1 Base models versus assistants See what pretraining alone produces and why it is not enough. Start
- 2 Instruction tuning and chat templates Understand supervised fine tuning on demonstrations and how roles are encoded. Start
- 3 Learning from preferences Learn how human comparisons become training signal through RLHF or direct methods. Start
- 4 Principles, AI feedback and verifiable rewards See newer ways to scale feedback and train reasoning. Start
- 5 System prompts at inference Understand how products steer a tuned model without retraining. Start
- 6 Side effects and trade offs Analyse the unintended behaviours post training can introduce. Start
Try asking
About this tutor
For technically minded learners who know the basics of language models and want to understand post training. You compare a raw pretrained model with a chat assistant, then walk through supervised fine tuning on demonstrations, reward modelling and reinforcement learning from human feedback, direct preference methods, AI feedback and written principles, safety training, and reinforcement learning on verifiable tasks. You also learn how chat templates, roles and system prompts work at inference time, and the side effects of post training such as sycophancy, over refusal and reward hacking. The focus is on concepts and trade offs, not a specific vendor's recipe.
Reviews
4.5
2 ratingsSample
- Tobias K.Sample
Framing each method as a fix for the previous one's problem made post training finally coherent. The side effects lesson was excellent.
- Leila F.Sample
Strong and careful about not claiming any lab's secret recipe. I wanted a little more depth on the KL penalty, but the intuition was clear.
About the teacher
I explain how language models really work, from tokens to attention, without hand waving
9 tutors 525 lessons taught Sample
I like taking the mystery out of language models. I teach what happens between typing a question and reading an answer: tokens, context windows, embeddings, attention, training and fine tuning. I start every topic with a picture or a small worked example and only add maths when it earns its place. My work background is in software and teaching workshops...
See Bastian's profile and tutorsMore like this
Other tutors on the same or nearby topics.