18,561 parameters (GRU: 15,425); 960 of them are the position table.
06-763 / L12
Attention and transformers, what it learned
06-763 / L12
Attention and transformers, the weakest prior
CV RMSE
best epoch
train–val gap
boosted trees
13.29
GRU
14.03
16
1.7
transformer
16.39
6
3.2
assumes almost nothing, so it must learn order and locality from about 70 engines
the same weak prior is why it wins at LLM scale
Zeng et al. 2023: one linear layer beat five forecasting transformers "in all cases" on nine datasets
06-763 / L12
Training a deep net well
06-763 / L12
Training a deep net well, the schedule
Minibatch gradient = true gradient + noise, and the step scales both. Near the minimum only the noise is left, so a constant rate never settles.
step: drop by a factor every few epochs
cosine: anneal smoothly to near zero (SGDR, CosineAnnealingLR)
one-cycle: up to a max and back down in one run (Smith)
06-763 / L12
Training a deep net well, the schedule on a noisy bowl
06-763 / L12
Training a deep net well, the schedule on our GRU
Five seeds, 60 epochs, stopping off; validation RMSE in cycles:
cosine
constant
best epoch, no stopping
13.01
12.91
early stopping, patience 10
13.04
13.04
last epoch (60)
14.55
15.54
The kept epoch is 9 to 17; with T_max = 60, cosine is still at 83% of its rate at epoch 17. Check the rate at the epoch you keep.
06-763 / L12
Training a deep net well, early stopping
Early stopping: stop when validation loss stops improving for patience epochs, and keep the best epoch's weights.
training error keeps falling; validation turns up once the net fits the training engines
seed 1: train 12.1 → 9.7, validation 13.4 → 16.1 from epoch 17 to 60
for a linear model under gradient descent it is L2 regularization (Goodfellow 7.8)
uses the validation split; the test set is still touched once
checkpointing: keep the best weights on disk, so a crash costs nothing
06-763 / L12
Training a deep net well, early stopping replayed
06-763 / L12
Training a deep net well, gradient clipping
one factor for the whole gradient: direction kept, length capped
clip_grad_norm_(model.parameters(), c) between backward() and step()
recurrent nets have steep walls in the loss; a step at a wall "would bring us very far" (Pascanu et al. 2013)
set c above ordinary norms; clip_grad_norm_ returns the norm, so log it
06-763 / L12
Training a deep net well, a question
clip_grad_norm_ with c = 1, and the gradient is g = (3, 4). What is g after clipping?
(1, 1)
(0.6, 0.8)
(0.43, 0.57)
(0.75, 1)
06-763 / L12
Training a deep net well, gradient clipping at a wall
06-763 / L12
Training a deep net well, weight decay
the ridge penalty from Lecture 9: shrink every weight by a fixed fraction each step
shrinks most where the data is weakest: factor per Hessian direction (Goodfellow 7.1.1)
Adam(weight_decay=...) is coupled L2; AdamW is true decay (Loshchilov and Hutter, Lecture 11)
06-763 / L12
Training a deep net well, weight decay shrinks the weak direction
06-763 / L12
Training a deep net well, reading the curves
loss curve
diagnosis
fix
both high, falling slowly
underfitting
more capacity, higher LR
train low, val rising
overfitting
regularize, stop earlier
oscillating or exploding
bad optimization
lower LR, normalize inputs
Lecture 11's broken loops were all the third kind, the one mistaken for a modeling problem.
06-763 / L12
Does the architecture earn its keep?
06-763 / L12
Does the architecture earn its keep?, the setup
Remaining useful life on C-MAPSS FD001 turbofans (from Lecture 8):
sliding 30-cycle windows over 17 sensor/setting channels
piecewise-linear RUL target (constant at 125, then linear)
GroupKFold by engine: no engine in both train and validation (Lecture 8)
Two models: a 1D-CNN on raw windows vs gradient boosting on window features (mean, std, last).
06-763 / L12
Does the architecture earn its keep?, the result
1D-CNN 19.9 ± 2.0 vs baseline 18.3 ± 1.0 cycles RMSE, four engine-grouped folds. The quick sequence model does not beat the baseline, and the gap is inside its own spread.
06-763 / L12
Does the architecture earn its keep?, the lesson
A deep net is not a default.
Grinsztajn et al. (45 datasets): "tree-based models remain state-of-the-art on medium-sized data" (~10K samples).
The CNN earns its advantage from scale, from tuning (schedule, regularization), and from raw structure a hand-crafted feature cannot capture. On a small, well-summarized benchmark, the baseline is the model to beat.
06-763 / L12
Limitations and trade-offs
06-763 / L12
Limitations and trade-offs
An architecture is an assumption to check. A CNN assumes locality and translation invariance; a global constraint or symmetry can make that a liability.
Depth needs data. A hundred engines is small; parameter sharing reduces but does not remove the appetite. Small data favors a strong classical model.
Accelerators and larger models have costs (Lecture 11: GPU 2.5x slower on a small model). Past the capacity the data supports, more layers buy overfitting. Debug on CPU.
06-763 / L12
Demo
l12-cnn-rul.ipynb
C-MAPSS FD001: 30-cycle sensor windows, piecewise-linear RUL, GroupKFold by engine. A small 1D-CNN in PyTorch, trained with a schedule and early stopping, per-epoch loss logged to MLflow, best checkpoint saved. Compared against the tabular baseline on the same grouped folds.
06-763 / L12
What to watch
The comparison, on a grouped split.
The sequence model does not automatically win.
The grouping by engine is what keeps the comparison honest instead of flattering.
06-763 / L12
Recap
Match architecture to input structure: vector to MLP, field to 2D-CNN, sequence to 1D-CNN or RNN
CNN economy = sparse interactions + parameter sharing; the receptive field grows with depth
Temporal CNN is the sensible default over an RNN for fixed windows
Train well: LR schedule, early stopping, clipping, weight decay, read the loss curves
A deep net is not a default; run a strong baseline, grouped by unit
06-763 / L12
Standings
Nicknames only. Everyone who skipped one still counted in every bar you saw.
06-763 / L12
Next
Reading PyTorch CNN/LSTM tutorials; Goodfellow ch. 9-10; Grinsztajn et al.
Full notes, with all sources: lectures/l12/notes.md
Play: the box walks the field, one output per position. Step twice and read the sum aloud. Switch to horizontal edge, then blur. Same 9 weights everywhere: 9 parameters for the map.
Start at stride 2, padding 1: 4x4. Set padding 0: 3x3, edge cells never centered. Stride 1, padding 1: 8x8, the 'same' convolution. Formula in the readout.
One input channel, three kernels, three output maps. Next layer's kernel is 3x3x3. Engineering inputs: u, v, p, T from CFD are 4 input channels, the same way RGB is 3.
Max vs average. No weights. 8x8 to 4x4; a 3x3 kernel after pooling spans 6x6 of the original, which is how deep layers get a wide view.
Expect A from MLP intuition. The engineering payoff: train on small, cheap simulation crops and apply to the full domain. The caveat in the why is real; a fully convolutional net is what keeps this property end to end.
Step with the slope kernel: output climbs toward failure. Switch to average, then random. Same k weights at every position.
3 layers x k=5 = 13 cycles. Last output: 7 real cycles, 6 padding zeros. Toggle dilation: 29 cycles, same weights.
Play at z=0.2, then drag z to 1 and to 0.05. Memory ~ 1/z. A GRU learns z per step from the input.
0.8^29 = 1.5e-3. B is the interesting wrong answer: it makes the cycle-1 gradient bigger in absolute terms but leaves the ratio alone, so early cycles still have no say relative to late ones.
Click through MLP, CNN, GRU, Transformer; hover the last hidden unit in each.
Step the query. Raise a: weights favour similar values. Toggle "shuffled" with no positional encoding: pooled summary identical. Turn positions on, shuffle: it changes.
The self-attention widget two slides back shows this with the shuffle toggle. The table on the next slide is the reveal. Ask whoever said B what would have to change for it to be exactly equal.
experiments.py order, seed 0. Reversed GRU is worse than predicting the mean (41.8).
Layer 1: entropy >= 99% of uniform. Layer 2, near-failure engine: last cycle gives 28% to last 5 cycles vs 17% uniform. Nearly uniform + mean pooling ~ a function of window averages, which boosting gets as features.
Test set: transformer 16.28, GRU 14.00. Last-cycle readout: 16.16, no help.
Same noise for both runs. At 0.12: last-30-step loss 0.030 cosine vs 0.133 constant. Drag the rate above 0.25: constant diverges, the stability limit along the steep axis is 2/8. Toggle step decay.
Honest null result. The schedule acts in epochs early stopping throws away. Seed spread is 1.8 cycles, so 13.01 vs 12.91 is a tie.
Play from epoch 1 and watch the counter. Seed 4, constant: patience 10 stops at 27 (13.74); the run reaches 13.06 at epoch 48, which needs patience 24. Is that real on 15 validation engines? Toggle cosine: same stop.
A is the common one and it is a real PyTorch function, clip_grad_value_, which is why it is worth separating the two out loud. C divides by 3 + 4 and D by the largest component: both keep the direction but use the wrong norm.
Unclipped: one step on the wall throws w from 1.5 to 11.5, loss 14.1, worse than the start. Right panel is our GRU: only epoch 1 had batches above 1.0 (a quarter); after that the max is 0.54. Insurance here.
lambda 0.5: optimum moves (2.40,1.60) -> (2.06,0.53). Firm weight keeps 86%, loose keeps 33%. Drag lambda: the green curve is the optimum for every lambda.