Advanced34 min
Mechanistic Interpretability
Features, circuits and superposition; induction heads, sparse autoencoders and activation patching as tools for finding what a network computes.
Not yet written
This lesson is on the syllabus but has no text yet
The full curriculum is published up front so you can see the whole route and its dependencies. Lessons are being written in curriculum order.
What it will cover
- circuits
- superposition
- SAE
- activation patching
- induction head
- probing