Reinforcement learning basics
Understand agents, rewards and policies, and compute a Q learning update yourself
A taste of a lesson
In the Q learning update, why not just replace the old value with the new target?
Because one experience can be misleading. Rewards and next states are often random, so a single lucky or unlucky step would throw the estimate around. The learning rate alpha blends old and new: with alpha 0.5 you move halfway towards the target each time, so the value settles towards an average over many experiences. Setting alpha to 1 means full replacement, which only works when everything is predictable. Tiny exercise: your Q value is 3, the target is 5 and alpha is 0.2. What is the new Q value?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Describe agent, environment, state, action, reward, policy and value in plain words
- Compute discounted returns and a Q learning update by hand
- Explain the exploration versus exploitation trade off using epsilon greedy
- Recognise reward hacking and why reward design is difficult
- Explain in broad terms how RL ideas are used to tune language models
Lesson plan
- 1 Agents, environments and rewards Name the parts of an RL problem using a small grid world. Start
- 2 Returns and discounting Compute discounted returns and see how the discount shapes behaviour. Start
- 3 Exploration versus exploitation Balance trying new actions against using the best known one in a bandit problem. Start
- 4 Values and Q learning Estimate action values and update them from experience with the Q learning rule. Start
- 5 Policies learned directly Understand policy gradient and actor critic methods in plain words. Start
- 6 Reward design and feedback from people See why rewards get gamed and how preference feedback is used for language models. Start
Try asking
About this tutor
For beginners who want to know how machines learn from trial and error rather than from labelled examples. You will meet the core ideas through small games: an agent, an environment, states, actions and rewards, and the goal of maximising total reward over time. You will compute discounted returns, see the exploration versus exploitation dilemma in a slot machine problem, and work through a Q learning update by hand. Later lessons explain policy gradient methods in words, why reward design is hard, and how reinforcement learning ideas appear when language models are tuned with human feedback. Light maths, all of it explained step by step.
Reviews
4.7
3 ratingsSample
- Sanna V.Sample
Free and genuinely good. The reward hacking examples were funny and made me think about badly designed targets at my own job.
- Lucia P.Sample
The slot machine lesson made exploration click instantly. I kept picking the first machine that paid out, which was exactly the point.
- Ahmed R.Sample
Clear and friendly. The Q learning numbers were helpful. I wanted more on how chatbots use this, but the tutor was honest that it is only a light overview here.
About the teacher
Architectures explained from the inside: convolutions, recurrence, attention and beyond
9 tutors 362 lessons taught Sample
I teach neural network architectures and the reasoning behind them. My working life has been spent implementing models from papers, getting them to train, and finding out which details the paper forgot to mention. I like to explain an architecture by asking what problem it was built to solve and what it costs, so convolutional networks, recurrent networks, transformers and...
See Nikolai's profile and tutorsMore like this
Other tutors on the same or nearby topics.