

Yeh (1998), UCI, CC BY 4.0 / NRMCA on concrete and strength tests

Rieth et al. (2017), CC0




Plus one synthetic set, where the truth is known: y = x1/3 + noise, for neural networks

Peng, Jury, Dönnes and Ciurtin (2021), Front. Pharmacol., CC BY 4.0
Classification predicts discrete categories (faulty/not faulty).
Regression predicts continuous values (temperature, pressure, flow rates).
| Term | Also called | What it is |
|---|---|---|
| Sample | Observation, row | One data point |
| Feature | Input, |
What the model is given |
| Target | Output, label, |
What the model predicts |
| Model | The function from features to target | |
| Training set | Data used to fit the parameters | |
| Test set | Data held back to check the fit | |
| Prediction | The model's output for new features | |
| Residual | Error | Actual minus predicted |




Fitting the model is one step of four:

The training set fits each model. The validation set compares models and chooses hyperparameters. The test set is used once, at the end, on the one model you chose.
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42,
)
| Metric | Name | Formula | What it tells you |
|---|---|---|---|
| Coefficient of determination | Share of the variance explained; 1 is perfect, 0 is the mean | ||
| MSE | Mean squared error | Punishes large errors; squared units | |
| RMSE | Root mean squared error | Same penalty, in the units of the target | |
| MAE | Mean absolute error | Robust to outliers; same units |
.fit on a linear model solves least squares (Lecture 7's lstsq); other models solve other problems behind the same .fit.
models = {
"linear": LinearRegression(),
"tree": DecisionTreeRegressor(),
"neural network": MLPRegressor(), # multi-layer perceptron
"Gaussian process": GaussianProcessRegressor(),
}
for name, model in models.items():
model.fit(X_train, y_train) # each learns its own parameters
y_pred = model.predict(X_valid) # validation rows, split off X_train
Four separate models, fitted and scored the same way so we can compare them; nothing is averaged or combined.
No constraints: the parameters are free. Most training problems are unconstrained; the Gaussian process will add bounds.
| Method | Gradient |
Step shaping |
|---|---|---|
| Gradient descent | All the rows | The identity |
| Stochastic gradient descent (SGD) | A random mini‑batch | The identity |
| Adam | A mini‑batch, averaged over steps (momentum) | Diagonal: a step size per parameter, from the running average of |
| L-BFGS | All the rows | Inverse-Hessian estimate from the last |
MLPRegressor defaults to Adam; solver="lbfgs" for small data; every Gaussian process (GP) fit is L-BFGS-B, L-BFGS with boundsLiu and Nocedal (1989) / Kingma and Ba (2015) / Bottou, Curtis and Nocedal (2018)
A straight line on the water data, one start: three optimizers reach the minimum, SGD bounces around it. L-BFGS takes 6 steps.
Water's pressure curve is not a straight line in
X = np.array([T**3, T**2, T, T**0]).T # columns: T³, T², T, 1
.fit is still least squares.
Feature engineering: transforming raw inputs into a form that makes the relationship easier for the model to capture.

Data: NIST Chemistry WebBook
Many candidate features, few rows: the curve chases the noise. Add a penalty to the loss:
Regularization: a penalty on the coefficients, added to the loss; its weight α sets how much they cost.
15 noisy points of
| Family | The idea | Good at | Bad at |
|---|---|---|---|
| Linear, with features | A weighted sum of chosen features | Little data, known physics | Shapes its features miss |
| Decision tree | Yes/no splits, a constant per region | Nonlinear, interacting effects; no scaling | Smooth functions, memorizing |
| Neural network | Layers of weights and activations | Any shape, large data | Small data, scaling, tuning |
| Gaussian process | A distribution over functions | Small data, an uncertainty |
No free lunch (Wolpert 1996): averaged over every possible problem, every method has the same error on new inputs. Our reading: a family wins when its assumptions match your problem, so compare families on held-out data.
A decision tree splits the data into regions with yes/no questions on the inputs, and predicts one constant in each region.

How a tree is grown: at each node, try every input and threshold, keep the split that lowers the squared error most, and repeat in each half. Each split is chosen once, never revisited: a greedy search.
Just as we expanded features with polynomials, a network expands the feature space adaptively, by learning nonlinear transformations.
Trivia: the 2024 Nobel Prize in Physics went to Hopfield and Hinton for foundational work on artificial neural networks; Hinton was on the CMU faculty from 1982 to 1987 (CMU News).
Noisy data from a true function,

Try a function with three nonlinear units, and fit its 10 parameters by curve fitting (optimization) on 80% of the points:
A neural network is layers of: multiply by weights, add biases, apply a nonlinear activation function. Training adjusts the weights and biases.

Each unit adds one tanh-shaped piece: unit 3 makes the steep early rise, unit 1 a steady slope, unit 2 a small kink at the end. The output adds them and

Neuron3.png, Egm4313.s12 (Prof. Loc Vu-Quoc), CC BY-SA 3.0
A ReLU unit,
Choices made before training, not fitted by it: the activation, the number of layers, the units per layer

MLPRegressor(
hidden_layer_sizes=(3,),
activation="tanh",
solver="lbfgs",
alpha=0.0,
max_iter=5000,
random_state=0,
)

The parameters minimize found are the weights and biases of a one-hidden-layer network: MLPRegressor scores

Ten starts: training sum of squared errors (SSE) 0.0796 to 0.1177, test random_state.
Two inputs that both matter, on very different scales:

Scale your inputs. The same network on the same data scores
A Gaussian process: a probability distribution over functions,
Rasmussen and Williams (2006), section 2.2
Each new point pulls the mean toward it and shrinks the band near it; far from the data the band stays wide

A Visual Exploration of Gaussian Processes (Distill) / Duvenaud, the Kernel Cookbook

RBF: length scale 0.235 (standardized, 12 training points), test
16 points from a GP design of experiments (Kitchin group), after Rehage and Hoffmann (1988)
Neural network
No constraints: the weights and biases are free
Gaussian process, in scikit-learn
lbfgs runs the same routine with no boundsThe surfactant data, with the signal and noise variances fixed at their fitted values; only the length scale
| Model | Variables | Objective | Kind of problem | Solved by |
|---|---|---|---|---|
| Linear | Squared error | Convex, one minimum | One least squares solve | |
| Ridge, lasso | + |
Convex | A solve; coordinate descent | |
| Decision tree | The splits | Squared error | Combinatorial | Greedy, split by split |
| Neural network | Squared error | Non-convex | L-BFGS, Adam, SGD | |
| Gaussian process | Non-convex, bounded | L-BFGS-B, restarts |
Start from the training problem, and keep adding terms:
Lecture 8 forecast the reactor pressure

Lecture 8's
Trained on 300 fault-free runs, tested on 100 others (Lecture 8's split by run):
| Model | Test RMSE (kPa) |
|---|---|
| Baseline: repeat the last value | 5.82 |
| Baseline: predict the mean | 7.57 |
| ARX, ridge (Lecture 8) | 4.71 |
| NN-NARX, 32 tanh units | 4.73 |
| GP-NARX, on 1,000 rows | 4.76 |

k-fold cross-validation: split the training data into

cv = KFold(
n_splits=5,
shuffle=True,
random_state=0,
)
scores = cross_val_score(
model, X_train, y_train,
cv=cv,
scoring="neg_root_mean_squared_error",
)
rmse = -scores.mean()

5-fold cross-validation on the 835 rows, rows assigned to folds at random:
| Model | RMSE (MPa) | MAE (MPa) | |
|---|---|---|---|
| Baseline: predict the mean | 16.8 | 13.6 | −0.01 |
| Linear | 10.6 | 8.46 | 0.60 |
| Linear, with physics features | 7.25 | 5.60 | 0.81 |
| Decision tree (no depth limit) | 6.92 | 4.52 | 0.83 |
| Neural network (16 tanh) | 5.96 | 4.20 | 0.87 |
| Gaussian process | 5.81 | 3.96 | 0.88 |
Physics features: log(age), and the water/cement ratio

Grouped cross-validation (GroupKFold): all the rows of a group (a mix, a batch, a run) are either all in training or all in validation.

Capacity: the range of functions a model can represent. More capacity fits more complicated relationships, and more of the noise.
Bias-variance trade-off: more capacity lowers bias and raises variance, so validation error is lowest in between.
Validation curve: training and validation error against one capacity knob, with the data fixed.

Learning curve: training and validation error against training-set size, model fixed.

| The curves show | Diagnosis | What helps | Here |
|---|---|---|---|
| Gap closed, error still high | Underfitting (high bias) | More capacity, better features | The line (left) |
| Gap closed, error low | A good fit | Stop, and test once | |
| Big gap: training low, validation much higher | Overfitting (high variance) | Less capacity, regularization | The tree (right) |
| Validation still falling at the right edge | Limited by data | More training samples |

All four refitted on the 21 points. At 300 °C, NIST: 517.7 MPa. Polynomial 192, tree flat at 100.7, network saturates at 148, GP 162 ± 232: the band still misses.
Chosen by hand today: 16 hidden units, a tree with no depth limit, the ridge
How do you search over
l09-regression.ipynbThe workflow on concrete, one step per cell: lock the test mixes, fit four families, compare KFold with GroupKFold, test once. Then NARX on Lecture 8's table.
Run it after class, top to bottom. The first run downloads the data; the cross-validation cell takes a minute or two.
Four questions to ask of any model, today's or your own:

Only its data.
The polynomial scored R2 = 0.9999975 on the water data, and gave 223 MPa at 300 °C, where NIST gives 517.7.

An optimization problem, and the family picks it.
A line has one minimum and a network many; a GP tunes a few bounded hyperparameters.

The question your split asked.
The tree beat the line on random folds and lost to it by 2 MPa on grouped folds.

A gap is variance, a high level is bias.
The tree kept a gap of 8.5 MPa. The line closed its gap but ended 1.4 MPa above a model with more capacity.
Start simple, and put what you know into the features: a line with two physics features came within 0.3 MPa of a Gaussian process on grouped folds.
Practice module for this session, for participation credit
Worked example l09-regression.ipynb, to run after class
110 minutes: lecture 90, questions 20. The notebook is a worked example students run after class; it is not shown live. Plan, by slide: opening and the examples (2-7) 9, types (9-12) 5, workflow (14-17) 6, training as optimization (19-23) 9, regression families (25-52) 36, cross-validation (54-59) 7, capacity (61-63) 5, limitations and the close (65-69) 5. About 82 minutes at a steady pace. If it runs long, the first to go, in order: the optimizer paths (23, one sentence), what each training solves (48), the kernel and prediction slide (43, the notes carry it), the NARX forecast figure (52).
A pause for the comic. "Just stir the pile until they start looking right" is the failure this session is about. ML: building models that learn patterns from data; instead of explicitly programming rules of physics/nature, we train algorithms to generalize from examples.
"Constant density" is the physics: a fixed mass of liquid water in a sealed, rigid container, so its density stays at 1000 kg/m3. It cannot expand when heated, so the pressure climbs steeply, from about 0.4 MPa near 0 C to about 100 MPa (roughly 1000 atm) at 100 C. First of today's examples; it comes back for feature engineering, regularization, the first tree, and at the end for what every family does outside its data.
The paper's system: the surfactant cetylpyridinium chloride (also the antiseptic in some over-the-counter mouthwashes) with the salt sodium salicylate. The long worms are wormlike micelles, and they tangle like polymer chains. Zero-shear viscosity is the plateau the viscosity reaches as the shear rate goes to zero: the thickness of the liquid sitting still. Why the peak: the salt screens the charges on the surfactant heads, so small spherical micelles grow into long worms and the viscosity climbs; past the peak it falls, commonly attributed to the worms branching or shortening (which of the two is still debated, Ziserman et al. 2009). The Kitchin group page labels the axis salt concentration and gives no units. The 16 points follow the shape of the paper's curve (sharp peak, dip, smaller second peak, fall) on a compressed scale: about 400-fold here, about five decades in the paper's own curve (replotted in Berret 2004, at 100 mmol/L of surfactant).
Sand and gravel are the file's fine and coarse aggregate. Slag (a byproduct of iron making) and fly ash (from coal power plants) replace part of the cement; superplasticizer is an additive that lets the fresh mix flow with less water. In standard practice a test result is the average of at least two cylinders crushed at the same age (NRMCA CIP 35); the UCI files do not say how many cylinders each row averages. Crushing destroys a cylinder, so a line is several cylinders of one mix, one per age. Designs specify strength at 28 days, which is why 425 of the 1,030 rows are 28-day tests. Only 428 distinct mixes, and 182 of them were crushed at several ages, which matters for the folds in the cross-validation section. High-performance concrete: concrete meeting performance and uniformity requirements that conventional ingredients and practice cannot always achieve (ACI).
Supervised: labels (target outputs). Unsupervised: no labels, find hidden structure (clustering, dimensionality reduction). Reinforcement: actions, state updates, feedback. Today is all supervised.
"ML is a bit jargonized. Let's clarify most of the terms used before we move on."
Four cards: what is one row, what goes in, what comes out, which task.
1 Feature engineering: select or transform the inputs (polynomial features of T, or the lagged columns of a NARX table: Lecture 7's regressors are features). Anything fitted to data, such as a scaler, is fitted after the split, on the training rows only. 2 Data splitting: training set to fit, test set to evaluate on unseen data. 3 Model selection: choose the algorithms. 4 Model validation: prevent overfitting, with metrics. Then the test set, once.
scikit-learn testimonials (https://scikit-learn.org/stable/testimonials/testimonials.html): is it used in real applications? Yes. The next section's question: what does each .fit solve?
The labels and their arrows come in one at a time.
The labels and their arrows come in one at a time. Same shape as the previous slide, now with the model's parameters as the decision variables and the data as constants. Regularization adds a penalty term to this objective, later in the regression section.
Four terms, one at a time. This one update covers every method on the next slide; they differ only in where g comes from and what H is.
One update for all of them: step = -eta H g. Gradient descent: H = I and the full gradient. SGD: H = I and the gradient of a random mini-batch, a noisy estimate of the full one; cheap steps, and the step length has to shrink for the iterates to settle. Adam: SGD plus two running averages, of g (momentum) and of g squared (a step size per parameter), bias-corrected because both start at zero; defaults 0.001, 0.9, 0.999. L-BFGS: quasi-Newton; keeps the last m pairs (s = the step, y = the change in gradient) instead of a Hessian, builds H g from them, line search; needs an exact full gradient, so small and medium data. scikit-learn's MLPRegressor docs: adam "works pretty well on relatively large datasets (with thousands of training samples or more)"; "for small datasets, however, 'lbfgs' can converge faster and perform better". The GP: L-BFGS-B, with bounds on every hyperparameter.
One clock for all four, the iteration number, on a log scale, so it speeds up as it runs: L-BFGS's six steps fill the first three seconds and the last 1,600 iterations take about three. A dot drops out when its method reaches the minimum, and its legend entry turns bold. SGD is drawn at every step to 400, then every 4th. A least squares problem, so the loss is a quadratic bowl, but a long narrow one: intercept and slope trade off (condition number 37.7). L-BFGS learns the shape of the valley in its first steps. SGD (4 rows per step, fixed step) reaches the floor fast and keeps bouncing: every mini-batch points somewhere slightly different. Adam runs on the full gradient here, so only its update differs from gradient descent; its momentum overshoots (the loop), then it settles.
"Merely a polynomial model, so you should not use it for extrapolation." Fitted with a library that implies ML, but no physics in it. At -50 C: 45.6 MPa. At 300 C: 223 against 517.7. Within the data range, a reasonable estimation.
Large alpha: a simpler model, and a risk of underfitting. Small alpha: plain least squares.
Small alpha: the curve chases the noise near the edges, training RMSE low, test RMSE higher. Large alpha: the curve flattens, both RMSEs rise. Lasso: the nonzero count falls as alpha grows. The right panel is a validation curve.
Linear regression first, because it is the one they know; this is why the other three exist. "Every possible problem" means every input-output relationship, weighted equally, and "new inputs" means points outside the training set (Wolpert's off-training-set error). Strictly, the 1996 theorem is proved for losses like zero-one, where a guess is right or wrong. For squared error the companion paper finds an edge, but all of it comes from where a method's guesses sit in the output range; on average, the way a method uses its data adds nothing. The surprise in Wolpert's abstract: it holds even for cross-validation against "anti-cross-validation" (pick the model with the largest validation error). So the second sentence is our reading and goes beyond the theorem: trusting validation is itself an assumption, that real problems have structure (smoothness, boxes, the right features) and that held-out points resemble the ones you will predict.
"Something interesting is happening here." Each split is the boundary that gives the minimum MSE. Depth 2 means four leaves, four values, a staircase. Test R2 0.911. Trees capture nonlinearity but overfit if the depth is too large, so max_depth is crucial. Finding the optimal tree is NP-complete (Hyafil and Rivest 1976), which is why it is grown greedily: no gradient, no starting point. Laird group: linear model decision trees inside optimization (OMLT).
One hidden unit at a time lights up and its term appears beside it, at the same height, in its color; then the output bias; then the sum. The previous slide's equation, one row per unit. The input weight w0k and bias b0k sit inside the tanh, the output weight w1k outside it. The deep version, several hidden layers, is on the hyperparameters slide.
Each colored curve is one unit from the previous slide: w0k and b0k set where the curve bends and how sharply, w1k sets how tall it is and its sign. The fitted terms are large and nearly cancel (offsets near -37, -23 and +64), so each is drawn shifted to start at zero; the offsets and b1 fold into one constant. The sum is the fit.
Play or the slider moves x. A unit lights up when it is on. Each kink is one unit switching on or off, so a ReLU network is piecewise linear. Unit 3 is on over the whole range, so it only adds a straight line.
Gradients by backpropagation. The worked example silences a ConvergenceWarning from the concrete network: the optimizer hit its 5,000-iteration limit before declaring convergence.
With x2 in the millions, all 20 tanh units saturate at plus or minus 1 on every training row, the gradients vanish and the network predicts the training mean (0.128) everywhere. Standardizing fixes it: zero mean, unit variance, fit on training rows only.
The picture from my F25 GP slides, built in four steps. It starts on a distribution of numbers, N(0, 1). First, fifteen draws land one at a time, each one a number; the newest is red. One, at -2.33, falls outside 2 std, where about 5% of draws should (4.55%). Second, the same idea one level up: the arrow, then a distribution of functions, its mean (dashed) and a band of 2 std. On the left the same range turns gray: at every x this GP's value is N(0, 1), the bell on the left, so the band is that gray range repeated at every x. Third, five draws, each a whole function, traced one after another. The gold one leaves the band at both ends: the band covers about 95% of the values at each x, so a whole curve can still cross. Fourth, the definition. How fast the curves wiggle is the kernel's length scale (0.3 here), two slides on. Formally (Rasmussen and Williams, Definition 2.1): any finite set of the function's values is jointly Gaussian. Parametric (a network): fix theta and fit f(x; theta). Nonparametric (a GP): a distribution over functions, and every prediction comes with an uncertainty. Going back a step undoes it; going forward again replays it.
Play or the slider adds the points one at a time. At 0 it is the prior, every function the kernel allows: mean 0 and a band of plus or minus 4.12 (2 std) everywhere. Each point is a noisy sample of f(u) = sin(u) + log(u) - exp(-0.1 u^2); the newest has a red ring. The band pinches at each point and stays wide in the gaps: average half-width 4.12, then 2.22 after 2 points, 0.80 after 5 and 0.16 after 20. The RMSE of the mean falls unevenly: near 0.45 from 7 to 9 points, then 0.083 at 10, when point 10 (u = 0.76) fills the gap at the left edge. The kernel (signal std 2.06, length scale 2.14, noise std 0.112) was fitted once on all 20 points and held fixed, so only the data change. Bayesian inference conditions on the data: the mean is a weighted average of the observed outputs, and the variance is what the data have not pinned down.
The squared exponential (RBF) kernel. The mean is a weighted average of the known outputs, with weights set by similarity; the variance starts at the prior variance and drops by what the data explain, so it grows far from the data. Other kernels: Matern (rougher), periodic, sums and products.
Strengths: probabilistic predictions, interpretable kernels, automatic Occam's razor. Weaknesses: O(N^3), kernel choice matters, needs scaling.
Checked in the scikit-learn source (1.6 and 1.9): GaussianProcessRegressor calls scipy.optimize.minimize(method="L-BFGS-B", bounds=kernel.bounds). Theta and the bounds are log-transformed, default (1e-5, 1e5) for the length scale, the constant (signal variance) and the white-noise level. alpha goes on the diagonal and is not optimized. Restarts are drawn log-uniformly inside the bounds. MLPRegressor(solver="lbfgs") calls the same routine without bounds; adam and sgd apply none. GPflow: L-BFGS-B on softplus-transformed parameters. GPyTorch: Adam on raw parameters with a softplus positivity constraint.
The labels and their arrows come in one at a time. The marginal likelihood averages over every function the GP prior allows, so a flexible kernel spreads its probability over many possible datasets and gives little to the one measured: that is the log-determinant term. My F25 slide: "GP estimation balances data fit and complexity of the predicted model: Occam's razor is automatic".
Short l: the mean spikes through every point and falls back to the average between them; complexity penalty large. The function may change between neighboring points. The sum reaches its minimum near l = 0.31 (the fitted value on all 16 points; the 0.235 two slides back is in standardized units, on 12 points). Long l: smooth over the whole range; the curve cannot reach the peak and the misfit explodes.
Four bullets, one at a time. F is any residual known to be zero: a balance, a rate law, an ODE evaluated at collocation points z_j. One sentence is enough.
Same fit/predict, on a lag table instead of a table of experiments. The GP-NARX got only 1,000 rows because its cost grows as N cubed: all 144,300 training rows would need a 167 GB kernel matrix (the notes have a section on where scikit-learn stops). ARX on the same 1,000 rows: 4.76. Lecture 7's deviation variables are the same idea as scaling the target.
Random folds: the 28-day row of a mix trains and the 56-day row of the same mix validates, so the model is asked about a mix it has already seen. Lecture 8's split by run, without the time axis.
ESL eq. 7.9: Err(x0) = noise + Bias^2 + Var. Noise: no model removes it. Bias: the average model's distance from the truth. Variance: how much the model moves when the training set changes. A straight line on concrete is the high-bias end; a tree with no depth limit is the high-variance end.
Training error falls to 0.95 and never to zero: nine settings (same mix, same age) were crushed more than once with different results, one from 22.9 to 55.9 MPa at 7 days. Those replicates are the noise term. GroupKFold validation stops improving at depth 9 (9.10) and wobbles 9.2 to 9.6. "Looking only at the train score can be misleading!" Random folds keep rewarding depth because they reward memorizing.
Each panel shows two things. The gap between the curves is the variance: the line's closes to 0.3 MPa as the training set grows, so it no longer overfits, and more samples can buy at most 0.3. The level where the curves meet is bias plus noise. "High" needs a reference: gradient-boosted trees on the same features and grouped folds reach 6.1 MPa, better than the line on all five folds. So about 1.4 MPa of the line's 7.4 on validation is bias that more capacity removes; the rest includes the replicate noise from the last slide. The line underfits. The same closed gap at a level nothing beats would be a good fit: stop and test once. The tree: 0.95 on its training samples, 9.4 on validation, a gap of 8.5: overfitting. Its validation curve is flat from about 400 samples on, so more samples of the same kind do not close the gap; hence no example here for the "limited by data" row. The curves average ten random orderings of the training rows. learning_curve does not shuffle by default, and in file order the tree's curve showed a false late drop.
A GP's uncertainty describes distance from the data under the kernel's smoothness assumption. It says nothing about physics the data never showed it. The notes' four-families table has a row for this: follows its features, flat, saturates, back to the mean. 192 here, not the 223 of the opening slide: this polynomial is fitted on all 21 points, that one on 16.
Not run in class; the 20 minutes after the deck are for questions. The notebook: the concrete workflow and the NARX forecasts from today's slides, one step per cell, on the real data. The first run downloads the concrete file (125 kB) and the fault-free TEP file (25 MB).
Each card's picture is the slide where the class saw the answer. The four questions hold for any model the students fit this semester, so the deck ends on them instead of a list. The habits under them: lock the test set and touch it once, scale the inputs and the target, set random_state.