Resources
Primary textbook
An Introduction to Statistical Learning with Applications in Python (ISLP) James, Witten, Hastie, Tibshirani, Taylor
Available free online at statlearning.com
This is the main reference for the course. The Python edition includes code examples in scikit-learn and statsmodels.
Reading by lecture
| Lecture | Topic | ISLP chapters |
|---|---|---|
| 3 | Preprocessing pipelines | Ch. 2 (Statistical Learning) |
| 4 | Linear models | Ch. 3 (Linear Regression) |
| 5 | Classification | Ch. 4 (Classification) |
| 6 | Bias-variance tradeoff | Ch. 2.2 (Bias-Variance) |
| 7 | Model selection | Ch. 5 (Resampling), Ch. 6 (Linear Model Selection) |
| 9 | Unsupervised / PCA | Ch. 12.1—12.2 (PCA) |
| 10 | Ensembles | Ch. 8 (Tree-Based Methods) |
| 11 | SVM | Ch. 9 (Support Vector Machines) |
| 12 | Boosting | Ch. 8.2.3 (Boosting) |
| 13 | Clustering | Ch. 12.4 (Clustering) |
| 14 | Neural networks | Ch. 10 (Deep Learning) |
Supplementary resources
Online references
- scikit-learn User Guide — authoritative documentation for all models and pipelines used in this course
- Python Data Science Handbook — good introduction to pandas, numpy, and matplotlib
- Kaggle Learn — short practical courses on ML topics
Additional reading
- The Elements of Statistical Learning (Hastie, Tibshirani, Friedman) — a more mathematical treatment, free at web.stanford.edu/~hastie/ElemStatLearn/
- Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow (Geron) — practical and code-heavy, good project reference