MST0052 · Lecture 9 · Fall 2026
90 minutes · no labels · one workhorse method
01
(X, y) — minimise error against y. A score to chase.
X only. No loss function, no test accuracy.
Compress many correlated columns into few informative ones — today.
Group similar rows without defining "similar" — Lecture 13.
200 questions — maybe 5 underlying attitudes.
No segments defined — yet.
Most pixels move together.
The L1 rule · support, not headline
02
03
The pattern · memorise this
PCA learns its rotation from the data — inside a Pipeline it is refit on each CV fold. No leakage.
Pipeline
2 or 3 — human eyes are the constraint.
PCA(n_components=0.9) keeps just enough for 90%.
PCA(n_components=0.9)
The downstream CV score decides — L7's rule.
04
Pitfall 1 · A rotation cannot bend
Linear combinations — straight directions through the cloud.
Curves, rings, manifolds — distorted or hidden.
Pitfall 2 · Naming is a story you impose
Calling PC1 a "size" axis is defensible.
Naming is wishful thinking. Call it PC1.
Pitfall 3 · The one that bites projects
Before you reach for PCA
TruncatedSVD
05
load_digits() · 8×8 images · 80/20 stratified · random_state=42
Step 1 · Fit PCA, read the cumsum
Step 2 · Project to 2D
Step 3 · Let CV choose n_components
n_components is a hyperparameter like any other — L7's machinery, unchanged.
n_components
Step 4 · Read the numbers
n_components=40
06
Beyond linear · visualisation only
Beyond linear · the kernel trick
Sparse data · TF-IDF
Option · PCA(whiten=True)
Pitfall 3 · the construction
Diagnostic · reconstruction error