MST0052 · Lecture 13 · Fall 2026
90 minutes · no labels, no ground truth · the last new method
01
Compress correlated columns into components.
Collect similar rows into clusters.
Split into a fixed k of disjoint clusters.
A tree of nested merges — every k at once.
Density-based — DBSCAN — in the backups.
The L1 rule · support, not headline
02
Assign · Update · Repeat
Spreads the starting centroids. Most of the protection.
Ten fresh starts, keep the best inertia. Free insurance.
Roughly spherical, similar-sized clusters. k known up front.
Elongated shapes, rare segments, unknown k.
03
scikit-learn clusters; scipy.cluster.hierarchy draws — linkage(X, 'ward') + dendrogram(Z).
scipy.cluster.hierarchy
linkage(X, 'ward')
dendrogram(Z)
Reference — screenshot this
Fast at large n. Needs k up front. Random init. Scatter + silhouette.
O(n²) and up. The cut chooses k. Deterministic. The dendrogram.
Default: k-means with n_init=10. Switch when n is small and the dendrogram is the deliverable.
n_init=10
04
Cohesion vs separation. Higher better — the default.
Nearest-cluster similarity. Lower better.
Between- vs within-variance. Higher better.
The minimum bar · if your project clusters
random_state
Pitfall 1 · Naming is a story you impose
"Cluster 2 is high-alcohol, low-flavanoid" — defensible.
"Cluster 2 is the premium segment" — wishful thinking. Call it cluster 2.
Pitfall 2 · One number is not evidence
Pitfall 3 · KMeans in a pipeline replaces your features
05
load_wine() · scaled · all 178 rows · labels locked away until Step 3
Step 1 · Sweep k, read both diagnostics
0.285 is weak by the >0.5 rule of thumb. Hold that thought.
Step 2 · Single inits across five seeds
Step 3 · ARI against the hidden classes
Step 4 · Cluster distances as features
CV accuracy · mean ± std
06
That was the last new method — the toolbox is complete.
Beyond k-means · density
Beyond k-means · soft assignments
Scale · hundreds of thousands of rows
Distance metrics · metric=
L9 callback · the standard plot