SkillHack › Career guides › Interpretability Researcher
How to become an Interpretability Researcher
Features, circuits, probing, and mechanistic analysis. Explain what is actually happening inside a model.
- 5phases in the roadmap
- 15topics to work through
- 75graded practice questions
- Freeno payment, ever
The Interpretability Researcher roadmap
Each topic below carries a study note, real reading, and practice questions in the app. A topic unlocks when you pass the one before it, so the order is the path, so you are never guessing what to learn next.
Phase 1Interpretability foundations
Why interpretability
Behavioral evals only sample the inputs you thought to test.
Behavioral vs mechanistic
Behavioral analysis characterizes what a model does from inputs and outputs, cheaply and at scale, but cannot say why.
Limits of interpretability
Current interpretability is partial: feature sets can be wrong, methods miss mechanisms they were not built to catch, and no-detection on tested inputs is bounded evidence, not proof of a global…
Phase 2Probing and attribution
Probing classifiers
A probe's accuracy conflates information in the representation with the probe's own capacity to fit labels.
Feature attribution
Attribution methods assign importance to input features, but methods like integrated gradients compute importance relative to a baseline that defines 'absence of signal.' The baseline choice…
Saliency and its pitfalls
A crisp, sensible-looking saliency map is not automatically a faithful one.
Phase 3Interpretability mechanistic
Circuits
Rather than describe a whole network, circuit analysis finds the subgraph of specific components (attention heads, MLP neurons) and connections that implement a behavior, characterizes each…
Features and superposition
Neurons are often polysemantic, firing for several unrelated concepts.
Sparse autoencoders
Sparse autoencoders decompose a layer's activations into a dictionary of sparse, hopefully monosemantic features.
Phase 4Interpretability evaluation
Validating interpretations
A story that fits every example you inspected is unfalsified narrative, and collecting more confirmations you selected is confirmation bias.
Causal interventions
Correlated activation is not causation.
Faithfulness
A plausible-looking rationale or attention pattern may not reflect the computation that produced the output; plausibility is not faithfulness.
Phase 5Interpretability applying
Interpretability for safety
In a safety case, interpretability is one layer of evidence whose strength is surfacing problems, hidden capabilities or deceptive reasoning that behavioral tests miss, integrated with evals and…
Debugging models
When a model fails on a subpopulation via a suspected spurious cue, interpretability drives debugging by locating what the model actually keys on, forming a specific hypothesis, and testing it…
Communicating findings
Most interpretability results are partial.
Reading for this path
The primary sources behind the topics above, all free to read.
- Zoom In: An Introduction to Circuits (Distill)
- Anthropic: Core Views on AI Safety
- A Mathematical Framework for Transformer Circuits
- Neel Nanda: Mechanistic Interpretability
- The Engineer's Interpretability Sequence (overview)
- Interpretability Dreams (transformer-circuits.pub)
- Designing and Interpreting Probes with Control Tasks
- Probing Classifiers: Promises, Shortcomings, Advances
- Axiomatic Attribution for Deep Networks (Integrated Gradients)
- The Building Blocks of Interpretability (Distill)
- Sanity Checks for Saliency Maps
- The (Un)reliability of Saliency Methods
- Toy Models of Superposition (transformer-circuits.pub)
- Towards Monosemanticity (transformer-circuits.pub)
Coming from another job?
Most people on this path arrived from somewhere else: engineering, analysis, testing, product, support, compliance, design. Onboarding asks what you do today and builds a short starter run out of the gaps, so you begin from what you already know rather than from chapter one. See how it works.
Other AI career paths
Knowledge Engineer (RAG)
Knowledge bases, embeddings, vector search, and graph RAG. Ground model answers in the right source of truth.
LLMOps Engineer
Prompt and version management, evaluation gates, and cost and latency monitoring. Keep LLM applications reliable through every model and prompt change.
ML Engineer
Model selection, training and evaluation, experimentation, and statistical reasoning about the models you ship.
MLOps Engineer
Model CI/CD, registries, monitoring, and drift detection. Keep the path from training to production repeatable and observed.
Multimodal AI Engineer
Text, image, audio, and video together - fusion, cross-modal retrieval, and multimodal integration. Build AI that sees, hears, and reads at once.
NLP Engineer
Text pipelines, classification, extraction, and multilingual systems. Build the language layer of AI products.
Start the Interpretability Researcher path for free
The full study notes, the reading, and the practice questions behind every topic above are in the app. Answer one tonight and you have started.
No payment, no credit card, no CV. Sign in with Google.