Interpretability: looking inside neural networks
Learn how researchers try to understand what happens inside models, and what we still cannot see
A taste of a lesson
Why can't researchers just look at each neuron to see what it does?
Because neurons are usually polysemantic: one neuron might fire for legal text, a certain font style and mentions of birds. The superposition hypothesis explains why. Models need to represent far more features than they have dimensions, so they store features as nearly orthogonal directions spread across many neurons. Reading one neuron is like reading one pixel of a compressed image. Sparse autoencoders try to recover those underlying directions as cleaner features. Question: why would sparsity make learned features easier for humans to interpret?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain the goals of mechanistic interpretability and how it differs from behavioural evaluation
- Describe probing, attribution and activation patching and what each can show
- Explain superposition, polysemantic neurons and sparse autoencoder features
- Describe steering by intervening on activations
- Assess the limits of current interpretability methods for safety claims
Lesson plan
- 1 Why look inside? Understand the motivations and the residual stream view. Start
- 2 Probes and attribution Learn what simple decoding and attribution methods reveal. Start
- 3 Causal methods and circuits Locate components responsible for behaviours. Start
- 4 Superposition and polysemantic neurons Understand why individual neurons are hard to interpret. Start
- 5 Sparse autoencoders and steering Learn how dictionary learning extracts interpretable features. Start
- 6 Limits and the road to verification Assess what interpretability can and cannot support today. Start
Try asking
About this tutor
For technically minded learners, researchers in adjacent fields and policy readers who want to understand mechanistic interpretability. You learn why understanding model internals matters for safety, debugging and trust, and you study the main tools: probing classifiers, attribution methods, activation patching, circuits, the superposition hypothesis, sparse autoencoders and feature dictionaries, and steering by intervening on activations. You also examine limits: scale, faithfulness of explanations, and how far current methods are from verifying a model's behaviour. Concepts are explained with diagrams in words and light linear algebra, without code.
Reviews
4.5
2 ratingsSample
- Mikael S.Sample
The superposition explanation finally made the research papers readable for me. Very careful about what is established versus early.
- Yasmin D.Sample
Rigorous and demanding. The patching experiment design exercise was excellent. Steering lesson felt a bit short.
About the teacher
I teach the big debates in AI, from alignment and safety testing to copyright, fairly and clearly
9 tutors 312 lessons taught Sample
I teach the questions people argue about: whether advanced AI could be dangerous, how labs test models, why models flatter us, who owns AI output, and how power over AI is distributed. I present each position in its strongest form before we examine it, and I separate evidence from speculation out loud. My background is in research communication and running...
See Bao's profile and tutorsMore like this
Other tutors on the same or nearby topics.