SkillHack › Career guides › Data Scientist
How to become an Data Scientist
Experiment design, statistics, causal inference, and modeling. Turn data into decisions and explain why.
- 5phases in the roadmap
- 15topics to work through
- 17graded practice questions
- Freeno payment, ever
The Data Scientist roadmap
Each topic below carries a study note, real reading, and practice questions in the app. A topic unlocks when you pass the one before it, so the order is the path, so you are never guessing what to learn next.
Phase 1Data science foundations
Statistics and probability
Statistics is the grammar of data science: distributions, sampling, and probability let you reason about uncertainty rather than react to noise.
Experimental design
Good experiments isolate a single cause by controlling everything else, usually through randomization.
Framing a data problem
The highest-leverage step happens before any modeling: turning a vague business ask into a precise question with a defined outcome, unit of analysis, success metric, and the decision it will inform.
Phase 2Data science analysis
Exploratory data analysis
EDA is the disciplined first look: distributions, summary statistics, correlations, and plots that reveal skew, outliers, and relationships before you commit to a model.
Data cleaning
Real data arrives with missing values, duplicates, and inconsistent encodings.
Feature engineering
Features encode domain knowledge into inputs a model can use: transformations, aggregations, and encodings.
Phase 3Data science modeling
Model selection
Choosing a model means matching method to data type, size, and constraints, and starting from strong baselines.
Validation and cross-validation
Honest performance estimates require testing on data the model never trained on.
Model interpretation
Interpretation explains what a model learned so stakeholders can trust and act on it.
Phase 4Data science experimentation
A/B testing
A/B tests are randomized experiments run in production.
Causal inference
When randomization is impossible, causal inference asks what would have happened otherwise.
Statistical significance
A p-value measures how surprising the data would be if there were no effect; it is not the probability the effect is large or important.
Phase 5Communication and production
Data visualization
Visualization is a perceptual tool: the right chart makes a pattern instantly readable, the wrong one hides or distorts it.
Storytelling with data
Analysis creates value only when it changes a decision.
Deploying models
Shipping a model is the start, not the end.
Reading for this path
The primary sources behind the topics above, all free to read.
- Khan Academy: Statistics and probability
- Seeing Theory (interactive probability)
- StatQuest: Design of experiments
- Khan Academy: Study design
- Google: Introduction to ML problem framing
- pandas: getting started tutorials
- seaborn: statistical data visualization
- scikit-learn: imputation of missing values
- pandas: working with missing data
- scikit-learn: preprocessing data
- Google: Feature engineering
- scikit-learn: choosing the right estimator
- scikit-learn: cross-validation
- scikit-learn: time series splitting
Coming from another job?
Most people on this path arrived from somewhere else: engineering, analysis, testing, product, support, compliance, design. Onboarding asks what you do today and builds a short starter run out of the gaps, so you begin from what you already know rather than from chapter one. See how it works.
Other AI career paths
Edge AI Engineer
On-device models, quantization for mobile and embedded, and offline inference. Run AI where the cloud can't reach.
Forward-Deployed Engineer
Workflow discovery, ambiguous requirements, agent-solution scoping, and deployment planning. Turn vague customer asks into scoped, shippable agent solutions.
Generative Media Engineer
Image, video, audio, and 3D generation - pipelines, controllability, and creative tooling. Build the systems behind generative media.
Inference Optimization Engineer
Quantization, distillation, serving performance, and latency and cost control. Make models fast and affordable at scale.
Interpretability Researcher
Features, circuits, probing, and mechanistic analysis. Explain what is actually happening inside a model.
Knowledge Engineer (RAG)
Knowledge bases, embeddings, vector search, and graph RAG. Ground model answers in the right source of truth.
Start the Data Scientist path for free
The full study notes, the reading, and the practice questions behind every topic above are in the app. Answer one tonight and you have started.
No payment, no credit card, no CV. Sign in with Google.