MST0052 · Lecture 5 · Fall 2026
90 minutes · three classifiers · one scoreboard
01
A continuous number — price, temperature.
A category — churn, spam, malignant.
"This customer will churn."
"This customer has a 73% chance of churning."
02
Baseline · same pipeline as L4
predict_proba is the output you evaluate with.
predict_proba
03
Memorises — follows noise. High variance.
Blurs — misses local patterns. High bias.
Tune k · same GridSearchCV as L4
Distances dominate by scale — StandardScaler is mandatory.
StandardScaler
Zero knobs · zero scaling
Text data? MultinomialNB — see the notes.
MultinomialNB
04
Rows = truth · Columns = prediction
Sick, flagged sick — caught it.
Sick, cleared — the dangerous cell.
Healthy, flagged — a false alarm.
Healthy, cleared — correctly.
scikit-learn prints negatives first: [[TN, FP], [FN, TP]].
[[TN, FP], [FN, TP]]
Reference — screenshot this
05
Predict malignant vs benign · binary
Step 1 · Load & split
Positive class = the one we must catch. stratify=y keeps the balance in both splits.
stratify=y
Step 2 · Three pipelines
Step 3 · Cross-validate
Step 4 · Compare
Step 5 · The confusion matrix, for real
06
Start with one parameter
Heavy imbalance · alternative to ROC
Three or more classes · automatic
When the probability itself matters
0.5 is a default, not a law