MST0052 · Lecture 12 · Fall 2026
90 minutes · the last new family · three families, one protocol
01
Independent bootstrap fits, averaged. Variance down, bias unchanged.
Sequential fits, each targets what's still wrong. Bias down too.
Once the variance is averaged away, there's nothing left to remove.
It goes after the residuals — and chases noise if the residuals are noise.
02
F ₘ = F ₘ₋₁ + η · hₘ — everything else is engineering.
03
Reference — screenshot this
They interact — lower η needs more trees; deeper trees want a smaller η.
Tiny corrections, many trees. Robust, slow.
Big corrections, few trees. Fast, overfits.
The pattern · L7 machinery, new model
04
Evaluate splits on every threshold — slow.
Split on bins — hundreds of times faster, same accuracy.
Why HistGradientBoosting, not the older GradientBoostingClassifier. Extra libraries: pip install — and commit your requirements.txt.
HistGradientBoosting
GradientBoostingClassifier
pip install
requirements.txt
Pitfall 1 · Six knobs, exploding grid
Pitfall 2 · Boosting wins by accident
Pitfall 3 · HistGBM has no feature_importances_
05
load_breast_cancer() · 80/20 stratified · random_state=42 · 5-fold CV · f1
Workflow 1 · No grid at all
Workflow 2 · learning_rate × max_depth, 12 cells
Point at the code · not run today
pip install xgboost — and commit your requirements.txt.
pip install xgboost
Test f1 · CV mean ± std underneath
Every CV gap is inside one fold-std — a three-way tie. The 114-row test set leans SVM.
CV f1 gaps · the knobs were forgiving
06
Linear models, trees, forests, SVMs, boosting — from here on, the craft is using them well.
L13 drops the labels: finding structure with no y at all.
History · reweighting, not gradients
Messy data · categorical_features=
Messy data · NaN just works
Interpretation · per-prediction
Search budgets · RandomizedSearchCV