Machine Learning Fundamentals Roadmap
Learn machine learning from scratch with this week-by-week roadmap: Python data stack, models, evaluation, and projects. (8 phases, 56 lessons, free).
Start This Roadmap Free →Phase 1: Python Foundations
- Python Environment Setup
Objectives: Explain what a virtual environment is and why it is needed for machine learning projects.; Create and activate a virtual environment using venv.; Install packages with pip inside the isolated environment.; De
- Data Types and Variables
Objectives: Understand the difference between mutable and immutable data types in Python.; Explain how variables store references to objects in memory.; Predict the outcome of operations on lists, tuples, and strings bas
- Control Flow Structures
Objectives: Explain what control flow structures are and why they are essential in programming.; Differentiate between conditional branching and loops, and when to use each.; Choose between for and while loops based on w
- Functions and Scope
Design pure functions with typed signatures and docstrings that transform tabular data, then trace variable resolution through local, enclosing, global, and built-in scopes.
- Data Structures for ML
Construct nested dictionaries and list comprehensions to represent a labeled dataset, then evaluate time-complexity trade-offs of lookup versus insertion operations.
- File I/O and Serialization
Parse CSV and JSON files into native structures using the standard library, then defend error-handling strategies for missing values and malformed records.
- Exploratory Data Script
Synthesize prior lessons into a single script that loads, cleans, summarizes, and visualizes a small CSV dataset with ASCII plots, producing a reusable template for Phase 2.
Phase 2: Data Exploration
- Data Types and Structures
Articulate the distinction between numerical, categorical, and ordinal data types and explain how their inherent structure dictates the choice of visualization and preprocessing technique.
- Descriptive Statistics
Derive and interpret measures of central tendency and dispersion from first principles to quantify the 'typical' value and spread of a dataset without relying on library black boxes.
- Distribution Shapes
Identify skewness, kurtosis, and multimodality in histograms and density plots to diagnose violations of normality assumptions before model selection.
- Correlation Analysis
Calculate Pearson and Spearman coefficients manually on a toy dataset to distinguish linear monotonic relationships from spurious correlations caused by outliers or non-linear dynamics.
- Missing Data Patterns
Classify missingness as MCAR, MAR, or MNAR by visualizing nullity matrices and correlation heatmaps to justify an imputation strategy rather than defaulting to mean filling.
- Outlier Detection
Apply the IQR method and Z-score thresholds on multivariate data to decide whether an extreme point represents noise, a data entry error, or a critical rare event worth preserving.
- Exploratory Dashboard
Synthesize profiling summaries, pair plots, and interactive widgets into a single reproducible notebook that communicates actionable insights and preprocessing decisions to a non-technical stakeholder.
Phase 3: Linear Regression
- Predicting with Lines
Articulate why a straight line is a reasonable first model for predicting a continuous target from one feature, using the analogy of drawing a trend line through a scatter plot.
- Residuals and Fit
Explain the concept of a residual as the vertical gap between a data point and the prediction line, and describe how the pattern of residuals reveals whether the line captures the underlying trend.
- Least Squares Criterion
Derive the sum of squared residuals as a single scalar loss function, and justify why squaring errors (rather than absolute values) leads to a unique, analytically solvable best-fit line.
- Closed Form Solution
Compute the optimal slope and intercept for simple linear regression using the normal equations, and interpret each term as a ratio of covariances to variances.
- Gradient Descent Mechanics
Implement batch gradient descent to minimize mean squared error, visualize the loss surface as a bowl, and articulate how the learning rate controls step size and convergence stability.
- Model Diagnostics
Evaluate fitted models using R-squared, residual plots, and hypothesis tests on coefficients, and diagnose violations of linearity, homoscedasticity, and independence assumptions.
- House Price Predictor
Build an end-to-end linear regression pipeline that loads a housing dataset, engineers a single predictive feature, trains via both closed-form and gradient descent, validates assumptions, and exports predictions for uns
Phase 4: Gradient Descent
- Optimization Landscape
Visualize the loss surface as terrain and articulate why gradient descent follows the steepest downhill direction to minimize error.
- Gradient Computation
Derive the gradient of mean squared error with respect to model parameters and explain how partial derivatives quantify parameter sensitivity.
- Learning Rate Dynamics
Experiment with learning rate values to observe convergence, oscillation, and divergence, then formulate guidelines for rate selection.
- Batch Gradient Descent
Implement batch gradient descent from scratch using vectorized operations and verify parameter updates match analytical expectations.
- Stochastic Gradient Descent
Contrast stochastic updates with batch updates by plotting loss trajectories and explaining the noise-convergence trade-off.
- Momentum Acceleration
Add momentum to the optimizer, visualize its dampening effect on oscillations, and justify why velocity terms escape shallow local minima.
- Linear Regression Trainer
Build a complete training pipeline that learns linear regression weights via gradient descent, logs loss curves, and exports the trained model for Phase 5.
Phase 5: Classification Basics
- Classification Concept
Articulate how classification differs from regression by examining discrete label assignment through visual decision boundary diagrams before formalizing the categorical prediction task.
- Logistic Regression Mechanics
Derive the sigmoid function from linear regression outputs to explain probability calibration, then implement binary classification from scratch using gradient descent to surface why linear models require non-linear acti
- Decision Boundary Visualization
Plot decision regions for logistic regression on synthetic 2D datasets to visually distinguish linear separability limits, then hypothesize how feature engineering or model complexity could resolve non-linear cases.
- Evaluation Metrics
Calculate accuracy, precision, recall, and F1-score from a confusion matrix by manually tracing predictions on an imbalanced dataset, then justify metric selection for specific business contexts like medical diagnosis ve
- K-Nearest Neighbors Classification
Implement KNN from scratch using Euclidean distance to classify points by majority vote, then experiment with k values to articulate the bias-variance tradeoff through visual boundary smoothness changes.
- Tree-Based Classification
Construct a decision tree by hand using entropy and information gain splits on a small dataset, then train sklearn's DecisionTreeClassifier to compare greedy splitting behavior against optimal boundaries.
- Classification Pipeline Build
Build an end-to-end classification system: preprocess a real dataset (e.g., breast cancer or iris), train logistic regression, KNN, and decision tree models, evaluate with cross-validation, and deploy the best model as a
Phase 6: Model Evaluation
- Generalization Gap
Articulate why a model's performance on training data diverges from its performance on unseen data by analyzing the tension between memorization and pattern extraction.
- Holdout Validation
Implement a train-test split protocol and justify the necessity of strict data separation to prevent information leakage during performance estimation.
- Cross Validation Strategy
Compare k-fold and stratified cross-validation approaches by reasoning through how each manages variance-bias tradeoffs in small-data regimes.
- Classification Metrics
Select and compute precision, recall, F1-score, and ROC-AUC for a binary classifier, defending each metric choice against specific business cost asymmetries.
- Regression Diagnostics
Interpret residual plots, MAE, RMSE, and R-squared to diagnose systematic prediction errors and quantify unexplained variance in continuous targets.
- Learning Curve Analysis
Diagnose high bias versus high variance by plotting training and validation error against dataset size, then prescribe data augmentation or model complexity adjustments.
- Evaluation Pipeline
Build a reusable evaluation module that automates cross-validated metric computation, statistical significance testing, and visual report generation for any scikit-learn estimator.
Phase 7: Decision Trees
- Decision Boundaries
Articulate how axis-aligned splits partition feature space by tracing classification regions on a 2D scatter plot.
- Impurity Measures
Compare Gini impurity and entropy on a toy dataset to explain why each quantifies node homogeneity differently.
- Recursive Splitting
Implement a greedy best-split search from scratch that recursively builds a tree until a stopping condition is met.
- Tree Depth Control
Visualize overfitting by plotting training versus validation accuracy across increasing max_depth values.
- Categorical Splits
Design a one-hot versus ordinal encoding experiment to show how split logic changes for nominal features.
- Missing Value Handling
Implement surrogate splits and compare their predictions against mean imputation on a dataset with 20% missing entries.
- Interpretable Classifier
Build a complete decision-tree pipeline — preprocessing, training, pruning, and rule extraction — that outputs human-readable if-then rules for a real-world dataset.
Phase 8: Ensemble Methods
- Wisdom of Crowds
Articulate why combining diverse predictions reduces error, using a visual analogy of averaged guesses converging on a true value.
- Bootstrap Aggregation
Explain how sampling with replacement and parallel training decorrelates base learners, then implement a bagged decision tree ensemble from scratch.
- Random Feature Subsets
Analyze how restricting split candidates at each node increases tree diversity, and extend the bagging implementation into a Random Forest.
- Sequential Error Correction
Contrast boosting's weighted re-training with bagging's independence, then build a simple AdaBoost loop that updates sample weights iteratively.
- Gradient Boosting Mechanics
Derive the pseudo-residual gradient for a differentiable loss function and implement a gradient boosting decision tree step-by-step.
- Stacking Meta-Learner
Design a stacking architecture where base model predictions become features for a meta-learner, and implement cross-validated prediction blending.
- Ensemble Model Showcase
Integrate bagging, boosting, and stacking into a single pipeline that outperforms any individual model on a structured tabular dataset.
Learn this with an AI mentor
Adaptive quizzes, weakness tracking, streaks — free.
Start This Roadmap Free →