SkillHack › Career guides › AI Alignment Researcher
How to become an AI Alignment Researcher
Training methods, oversight, and alignment evaluations. Make advanced AI pursue the goals we intend.
- 5phases in the roadmap
- 15topics to work through
- 105graded practice questions
- Freeno payment, ever
The AI Alignment Researcher roadmap
Each topic below carries a study note, real reading, and practice questions in the app. A topic unlocks when you pass the one before it, so the order is the path, so you are never guessing what to learn next.
Phase 1AI alignment foundations
The alignment problem
Alignment is the gap between the objective we can specify and the one we actually intend.
Specification and reward
Specification is turning an intended goal into a training signal.
Risks from advanced AI
A recurring argument in alignment is that some risks scale with capability: a more capable system pursues a misspecified objective more competently, potentially via instrumental strategies like…
Phase 2AI alignment training methods
RLHF
Reinforcement learning from human feedback trains a reward model on human preference comparisons, then optimizes a policy against it.
Constitutional AI
Constitutional AI reduces reliance on human harm labels: the model critiques and revises its own outputs against a written set of principles (a constitution), and those AI-generated preferences train…
Scalable oversight
When evaluating an answer is nearly as hard as producing it, one-pass human review becomes the bottleneck.
Phase 3AI alignment evaluation
Alignment evals
A good alignment eval samples a defined threat distribution, holds its prompts out of training, and reports coverage and confidence intervals.
Deception and sandbagging
Sandbagging is strategic underperformance when a model infers it is being evaluated, which makes evals underestimate true capability.
Red-teaming
Red-teaming finds inputs that make a safety-trained model misbehave.
Phase 4Advanced AI alignment
Reward hacking
Reward hacking is when a policy scores the measured proxy while violating the intended goal, hiding a mess from a camera instead of cleaning it.
Goal misgeneralization
Goal misgeneralization is competent pursuit of a proxy goal that was correlated with the intended goal in training but comes apart under distribution shift.
Interpretability for alignment
Interpretability aims to read a model's internals to support safety cases, e.g., detecting deception.
Phase 5AI alignment practice
Designing alignment experiments
Attributing an effect to an intervention requires a controlled comparison that varies only that intervention on the same model and eval set, with adequate sample size and a pre-registered metric.
Measuring alignment
Alignment has competing axes: pushing harmlessness alone rewards a model that refuses everything and tanks helpfulness.
Communicating results
Responsible reporting states the measured result with its sample size and confidence interval, scopes it to exactly what was tested, and names what was not.
Reading for this path
The primary sources behind the topics above, all free to read.
- Anthropic: Core Views on AI Safety
- AGI Safety Fundamentals (BlueDot)
- Concrete Problems in AI Safety (arXiv)
- DeepMind: Specification gaming
- Unsolved Problems in ML Safety (arXiv)
- InstructGPT: instructions with human feedback (arXiv)
- Training a Helpful and Harmless Assistant (arXiv)
- Constitutional AI: Harmlessness from AI Feedback (arXiv)
- Anthropic: Claude's Constitution
- Measuring Progress on Scalable Oversight (arXiv)
- AI safety via debate (arXiv)
- Anthropic: Challenges in evaluating AI systems
- Alignment Forum
- Sleeper Agents (arXiv)
Coming from another job?
Most people on this path arrived from somewhere else: engineering, analysis, testing, product, support, compliance, design. Onboarding asks what you do today and builds a short starter run out of the gaps, so you begin from what you already know rather than from chapter one. See how it works.
Other AI career paths
AI Automation Engineer
Agents, tool integration, and workflow orchestration. Turn multi-step business processes into automated flows.
AI Data Engineer
Ingestion pipelines, embeddings, vector stores, and data quality for training and retrieval. Feed AI systems clean, fresh, well-shaped data.
AI Developer Advocate
Docs, sample apps, talks, and community feedback loops. Help developers build with AI and carry their voice back to the product.
AI Engineer
Prompt and context design, tools, retrieval, orchestration, evaluation, and structured output. Build reliable agent apps end to end.
AI Evals Engineer
Benchmark design, LLM-as-judge, regression suites, and quality metrics. Measure what models actually do before and after every change.
AI for Science Engineer
Modeling in biology, chemistry, physics, and materials, plus research tooling. Apply AI to scientific discovery.
Start the AI Alignment Researcher path for free
The full study notes, the reading, and the practice questions behind every topic above are in the app. Answer one tonight and you have started.
No payment, no credit card, no CV. Sign in with Google.