MST0052 · Lecture 14 · Fall 2026
90 minutes · no new tool · vocabulary + judgment
01
Validation, leakage, comparison, interpretation — defensible at the oral.
Often hides exactly those decisions behind library APIs.
What a network is, in terms you already own.
Epoch, batch, backprop, attention — enough to read on.
When classical wins — for the oral, and the next decade.
02
Translation · same math, new words
β · β₀ · σ · log-likelihood · fitting
weights w · bias b · activation · cross-entropy loss · training
Reference — screenshot this
03
Squared error — L4's loss.
Cross-entropy — the log-loss from L5.
Exact gradient, slow steps.
Noisy gradient, many more steps — and the noise mildly regularises.
04
Edges.
Motifs.
Objects.
Classical ML — structure is encoded in the model.
Deep learning starts to win.
Deep learning — the structure is perceptual.
05
A coefficient table.
Permutation importance.
A black box, plus post-hoc tools of uncertain faithfulness.
The heuristic · default — deviate only with a defence
load_breast_cancer() · 80/20 stratified · random_state=42 · 5-fold CV · f1
Step 1 · Two hidden layers, scaled, early stopping
Single fit: 0.02 s — speed was never the problem here.
Step 2 · Five seeds, same spec
Step 3 · early_stopping on 455 rows
Test f1 · CV mean ± std underneath
Forest (L10): test 0.958 · CV 0.972 ± 0.014. CV gaps sit near one fold-std — but the MLP needed a flag flipped to join the tie. Champion unchanged.
The counterfactual · raw biopsy images
06
The toolbox is closed — L15 is the rehearsal workshop. Bring your project.
Beyond sklearn · the same model, five lines
Architectures · images
Architectures · attention
Translation · the L6 dial in new clothes
Foundation models · the decision