06-763 / L9

Lecture 9: The machine learning workflow II, regression

Week 5, Machine learning and deep learning

Systems and Toolchains for AI Engineers

Systems and Toolchains for AI Engineers
06-763 / L9

What today is about

Lecture 8: a time series, where order in time decided everything. Today: mostly rows with no time order, one experiment each.

  1. Four examples, and the types of machine learning
  2. The workflow, and training as optimization
  3. Regression: linear, trees, neural networks, Gaussian processes, NARX (nonlinear)
  4. Cross-validation and model capacity
  5. Limitations, and one question left open (hyperparameters)
  6. A worked example, to run after class

xkcd 1838, Randall Munroe, CC BY-NC 2.5

Systems and Toolchains for AI Engineers
06-763 / L9

Why this matters, what stirring the pile looks like

Water sealed in a rigid container: how fast does its pressure rise as it heats? Data from NIST (National Institute of Standards and Technology)

  • Fit a third-degree polynomial, , to 16 of the 21 points
  • Training = 0.9999975
  • At 300 °C it says 223 MPa; NIST says 517.7 MPa
  • The score only measured the model where the data is

NIST source.

Systems and Toolchains for AI Engineers
06-763 / L9

Today's examples, surfactant viscosity

  • Surfactant: a soap-like molecule; with salt, many join into long, tangled worms
  • Zero-shear viscosity: how thick the liquid is at rest, in the bottle or your hand
  • Salt is the knob: a little thickens shampoo, too much makes it runny again
  • In our data, viscosity peaks with salt addition, then falls

Rehage and Hoffmann (1988) / Kitchin group

Systems and Toolchains for AI Engineers
06-763 / L9

Today's examples, concrete strength

  • Concrete: cement and water glue sand and gravel together; it keeps hardening for months
  • A mix: one recipe, kg per m³ of 7 ingredients: cement, water, sand, gravel, slag, fly ash, superplasticizer
  • A test: cast cylinders from a mix, crush them days later; strength is the breaking stress (MPa)
  • Each gray dot: one test, at its age and strength; one row of the data (1,030 rows)
  • Each colored line: one mix, its cylinders crushed at several ages, 3 days to a year

Yeh (1998), UCI, CC BY 4.0 / NRMCA on concrete and strength tests

Systems and Toolchains for AI Engineers
06-763 / L9

Today's examples, the Tennessee Eastman process (TEP)

  • Lecture 8's simulated plant: 52 channels; each row is one 3-minute sample of a run, so rows are split by run
  • Today's job: forecast the reactor pressure 30 minutes ahead (NARX, Lecture 7's model on lagged outputs and inputs)

Rieth et al. (2017), CC0

Systems and Toolchains for AI Engineers
06-763 / L9

Today's examples

Water

  • NIST, 21 temperatures
  • Feature engineering, extrapolation

Surfactant

  • 16 experiments
  • Gaussian processes

Concrete

  • 1,030 rows, 428 mixes
  • Comparing models, cross-validation

Tennessee Eastman

  • Lecture 8's plant
  • NARX forecasts
Plus one synthetic set, where the truth is known: y = x1/3 + noise, for neural networks
Systems and Toolchains for AI Engineers
06-763 / L9

Types of machine learning

Systems and Toolchains for AI Engineers
06-763 / L9

Types of machine learning

Peng, Jury, Dönnes and Ciurtin (2021), Front. Pharmacol., CC BY 4.0

Systems and Toolchains for AI Engineers
06-763 / L9

Types of machine learning, two supervised tasks

Classification predicts discrete categories (faulty/not faulty).
Regression predicts continuous values (temperature, pressure, flow rates).

  • Most engineering problems: supervised regression
  • Process control: reinforcement learning and classification (fault diagnosis) are common too
Systems and Toolchains for AI Engineers
06-763 / L9

Types of machine learning, some terminology

Term Also called What it is
Sample Observation, row One data point
Feature Input, What the model is given
Target Output, label, What the model predicts
Model The function from features to target
Training set Data used to fit the parameters
Test set Data held back to check the fit
Prediction The model's output for new features
Residual Error Actual minus predicted
Systems and Toolchains for AI Engineers
06-763 / L9

Today's examples, in this vocabulary

Water

  • Sample: one temperature
  • Feature: temperature
  • Target: pressure
  • Task: regression

Surfactant

  • Sample: one solution
  • Feature: log concentration
  • Target: log viscosity
  • Task: regression

Concrete

  • Sample: one mix at one age
  • Features: 7 ingredients, age
  • Target: strength
  • Task: regression

Tennessee Eastman

  • Sample: one 3-minute step
  • Features: 10 past pressures, 11 valves
  • Target: pressure in 30 min
  • Task: regression (NARX)
Systems and Toolchains for AI Engineers
06-763 / L9

The machine learning workflow

Systems and Toolchains for AI Engineers
06-763 / L9

The workflow

Fitting the model is one step of four:

Systems and Toolchains for AI Engineers
06-763 / L9

The workflow, three sets

The training set fits each model. The validation set compares models and chooses hyperparameters. The test set is used once, at the end, on the one model you chose.

  • Look at the test score and change something: it has become validation data
  • The test number is the number you report, even when you do not like it
X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,
    random_state=42,
)
Systems and Toolchains for AI Engineers
06-763 / L9

The workflow, scoring a regression model

Metric Name Formula What it tells you
Coefficient of determination Share of the variance explained; 1 is perfect, 0 is the mean
MSE Mean squared error Punishes large errors; squared units
RMSE Root mean squared error Same penalty, in the units of the target
MAE Mean absolute error Robust to outliers; same units
  • RMSE much larger than MAE: a few large errors among many small ones
  • Parity plot: predicted against measured; the diagonal is the perfect model
Systems and Toolchains for AI Engineers
06-763 / L9

The workflow, where is the training here?

.fit on a linear model solves least squares (Lecture 7's lstsq); other models solve other problems behind the same .fit.

models = {
    "linear": LinearRegression(),
    "tree": DecisionTreeRegressor(),
    "neural network": MLPRegressor(),        # multi-layer perceptron
    "Gaussian process": GaussianProcessRegressor(),
}
for name, model in models.items():
    model.fit(X_train, y_train)        # each learns its own parameters
    y_pred = model.predict(X_valid)    # validation rows, split off X_train

Four separate models, fitted and scored the same way so we can compare them; nothing is averaged or combined.

Systems and Toolchains for AI Engineers
06-763 / L9

Training is an optimization problem

Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, the problem you already know

min z f(z) s.t. h(z) = 0 g(z) ≤ 0
  • Decision variables : flows, temperatures, a design
  • Objective : cost, energy, lost yield
  • Equality constraints : mass and energy balances, equilibrium
  • Inequality constraints : bounds, purity, safety limits
Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, a model is f(x; θ)

min θ L(θ) = Σ i (yi − f(xi; θ))2
  • Decision variables : the model's parameters (coefficients, weights)
  • Objective : the loss, how badly the model misses the training data
  • The model : turns the inputs into a prediction
  • The data : fixed numbers, the constants of the problem

No constraints: the parameters are free. Most training problems are unconstrained; the Gaussian process will add bounds.

Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, how an optimizer moves

θk+1 = θk − ηk Hk gk
  • Current point : where the parameters are now, at iteration
  • Step length : how far to move
  • Gradient : the uphill direction, so we step against it; computed from all the rows, or from a random mini-batch of a few rows
  • Step shaping : a matrix that rescales the step using the curvature of the loss; the identity matrix means no reshaping, plain gradient descent
Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, four methods

Method Gradient from Step shaping
Gradient descent All the rows The identity
Stochastic gradient descent (SGD) A random mini‑batch The identity
Adam A mini‑batch, averaged over steps (momentum) Diagonal: a step size per parameter, from the running average of
L-BFGS All the rows Inverse-Hessian estimate from the last steps, line search
  • In scikit-learn: MLPRegressor defaults to Adam; solver="lbfgs" for small data; every Gaussian process (GP) fit is L-BFGS-B, L-BFGS with bounds

Liu and Nocedal (1989) / Kingma and Ba (2015) / Bottou, Curtis and Nocedal (2018)

Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, four optimizers on one problem

Iteration 0

A straight line on the water data, one start: three optimizers reach the minimum, SGD bounces around it. L-BFGS takes 6 steps.

Systems and Toolchains for AI Engineers
06-763 / L9

Regression models

Systems and Toolchains for AI Engineers
06-763 / L9

Linear regression and feature engineering

, and "learn" the coefficients

Water's pressure curve is not a straight line in , so we give the model more columns:

X = np.array([T**3, T**2, T, T**0]).T    # columns: T³, T², T, 1

: a third-degree polynomial, still linear in the coefficients, so .fit is still least squares.

Feature engineering: transforming raw inputs into a form that makes the relationship easier for the model to capture.

Systems and Toolchains for AI Engineers
06-763 / L9

Linear regression, the water data

Data: NIST Chemistry WebBook

Systems and Toolchains for AI Engineers
06-763 / L9

Regularization, why

Many candidate features, few rows: the curve chases the noise. Add a penalty to the loss:

Regularization: a penalty on the coefficients, added to the loss; its weight α sets how much they cost.

  • Ridge, : shrinks weak coefficients, rarely to zero
  • Lasso, : sets weak coefficients to exactly zero
  • A penalty method: the penalty replaces the constraint (ESL 3.4.1)
Systems and Toolchains for AI Engineers
06-763 / L9

Regularization, live

15 noisy points of , fitted with a 12th-degree polynomial: 12 candidate features.

log10 α α = 1e-12
Systems and Toolchains for AI Engineers
06-763 / L9

Regression models, four families

Family The idea Good at Bad at
Linear, with features A weighted sum of chosen features Little data, known physics Shapes its features miss
Decision tree Yes/no splits, a constant per region Nonlinear, interacting effects; no scaling Smooth functions, memorizing
Neural network Layers of weights and activations Any shape, large data Small data, scaling, tuning
Gaussian process A distribution over functions Small data, an uncertainty , the kernel choice

No free lunch (Wolpert 1996): averaged over every possible problem, every method has the same error on new inputs. Our reading: a family wins when its assumptions match your problem, so compare families on held-out data.

Wolpert (1996) / Wolpert (2020), free overview

Systems and Toolchains for AI Engineers
06-763 / L9

Decision trees

A decision tree splits the data into regions with yes/no questions on the inputs, and predicts one constant in each region.

How a tree is grown: at each node, try every input and threshold, keep the split that lowers the squared error most, and repeat in each half. Each split is chosen once, never revisited: a greedy search.

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks

  • In chemical engineering we often face nonlinear models (kinetics, transport, thermodynamics), where linear and polynomial regression are limited
  • Neural networks emerged in the 1940s and 1950s, inspired by biological neurons, were revived in the 1980s, and became dominant in the 2010s
  • Today we use them as universal function approximators

Just as we expanded features with polynomials, a network expands the feature space adaptively, by learning nonlinear transformations.

Trivia: the 2024 Nobel Prize in Physics went to Hopfield and Hinton for foundational work on artificial neural networks; Hinton was on the CMU faculty from 1982 to 1987 (CMU News).

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, a flexible nonlinear regression

Noisy data from a true function, , with :

Try a function with three nonlinear units, and fit its 10 parameters by curve fitting (optimization) on 80% of the points:

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, the structure

A neural network is layers of: multiply by weights, add biases, apply a nonlinear activation function. Training adjusts the weights and biases.

Input Hidden layer (tanh) Output (linear) x y w00 w10 +b00 w01 w11 +b01 w02 w12 +b02 +b1
Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, how the terms add up

Each unit adds one tanh-shaped piece: unit 3 makes the steep early rise, unit 1 a steady slope, unit 2 a small kink at the end. The output adds them and ; the colors match the network diagram.

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, why "neural"

  • Neurons combine input signals and "fire" if strong enough; an activation function is the model of that
  • Far from real brains, but the language stuck

Neuron3.png, Egm4313.s12 (Prof. Loc Vu-Quoc), CC BY-SA 3.0

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, neurons firing

A ReLU unit, , fires once its weighted input passes zero. Five of them, fitted to plus noise:

x = 0 1
Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, hyperparameters

Choices made before training, not fitted by it: the activation, the number of layers, the units per layer

MLPRegressor(
    hidden_layer_sizes=(3,),
    activation="tanh",
    solver="lbfgs",
    alpha=0.0,
    max_iter=5000,
    random_state=0,
)
Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, the same model twice

The parameters minimize found are the weights and biases of a one-hidden-layer network: MLPRegressor scores = 0.950 on the same test points. Just math, not magic.

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, what training solves

: non-convex, local minima, sensitive to the start

Ten starts: training sum of squared errors (SSE) 0.0796 to 0.1177, test 0.920 to 0.950. Always set random_state.

Systems and Toolchains for AI Engineers
06-763 / L9

Neural networks, scaling!

Two inputs that both matter, on very different scales: and

Scale your inputs. The same network on the same data scores unscaled, worse than predicting the average, and once each input has mean 0 and standard deviation 1.

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes

A Gaussian process: a probability distribution over functions, , set by a mean function and a kernel (how similar two inputs are).

Rasmussen and Williams (2006), section 2.2

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, from prior to posterior

Prior 20 points

Each new point pulls the mean toward it and shrinks the band near it; far from the data the band stays wide

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, the kernel and the prediction

  • The length scale : how far apart two inputs can be and still be similar
  • Prediction at a new , with the similarity of the training points (noise included):

A Visual Exploration of Gaussian Processes (Distill) / Duvenaud, the Kernel Cookbook

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, surfactant viscosity

RBF: length scale 0.235 (standardized, 12 training points), test = 0.783 / Matérn: 0.836

16 points from a GP design of experiments (Kitchin group), after Rehage and Hoffmann (1988)

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, what training solves

Neural network

No constraints: the weights and biases are free

Gaussian process, in scikit-learn

: a few kernel hyperparameters, bounded

  • In scikit-learn the GP runs L-BFGS-B with those bounds; the network's lbfgs runs the same routine with no bounds
  • GPflow and GPyTorch keep the hyperparameters positive with a transform (softplus) and optimize freely
Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, the marginal likelihood

−log p(y | X, θ) = 1/2 yTK−1y + 1/2 log |K| + N/2 log 2π
  • What training minimizes: minus the log of the marginal likelihood, the probability of the measured data for these kernel settings
  • Misfit: large when the kernel explains the data poorly
  • Complexity penalty: large when the kernel is flexible enough to explain almost any data
  • A constant: is the number of points; it does not change training

is the kernel's similarity matrix, noise included. Minimizing the sum picks the simplest kernel that still explains the data: Occam's razor, built in.

Systems and Toolchains for AI Engineers
06-763 / L9

Gaussian processes, the length scale

The surfactant data, with the signal and noise variances fixed at their fitted values; only the length scale (in log concentration) moves.

Length scale: short Long
Systems and Toolchains for AI Engineers
06-763 / L9

Regression models, what each training solves

Model Variables Objective Kind of problem Solved by
Linear Squared error Convex, one minimum One least squares solve
Ridge, lasso + , + Convex A solve; coordinate descent
Decision tree The splits Squared error Combinatorial Greedy, split by split
Neural network Squared error Non-convex L-BFGS, Adam, SGD
Gaussian process Non-convex, bounded L-BFGS-B, restarts
Systems and Toolchains for AI Engineers
06-763 / L9

Training as optimization, with physics in it (spoiler alert! :) )

Start from the training problem, and keep adding terms:

  • Fit the data:
  • Add regularization, prefer small weights:
  • What if the penalty is a law the model must obey, a mass balance or an ODE?
  • We will see later in the course how to blend physics and machine learning: scientific machine learning

Raissi, Perdikaris and Karniadakis (2019)

Systems and Toolchains for AI Engineers
06-763 / L9

Regression models, back to Lecture 8: NARX

Lecture 8 forecast the reactor pressure 30 minutes ahead, from the plant's recent past:

Lecture 8's was linear (ridge): an ARX model. A network or a GP as makes it NARX, a nonlinear ARX.

Systems and Toolchains for AI Engineers
06-763 / L9

Regression models, NARX results

Trained on 300 fault-free runs, tested on 100 others (Lecture 8's split by run):

Model Test RMSE (kPa)
Baseline: repeat the last value 5.82
Baseline: predict the mean 7.57
ARX, ridge (Lecture 8) 4.71
NN-NARX, 32 tanh units 4.73
GP-NARX, on 1,000 rows 4.76
  • Scale the target too: pressure sits near 2,705 kPa with a spread of 7.7; an unscaled network just predicts the mean (7.57)
  • The three tie: held at its operating point, the plant behaves close to linearly
Systems and Toolchains for AI Engineers
06-763 / L9

Regression models, NARX forecasts

  • All three forecasts lie on top of each other; the GP adds a band, with 94.8% of test points inside 2 standard deviations
Systems and Toolchains for AI Engineers
06-763 / L9

Choosing a model with cross-validation

Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation

k-fold cross-validation: split the training data into folds, train times holding out a different fold each time, and average the scores.

Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation, in scikit-learn

cv = KFold(
    n_splits=5,
    shuffle=True,
    random_state=0,
)
scores = cross_val_score(
    model, X_train, y_train,
    cv=cv,
    scoring="neg_root_mean_squared_error",
)
rmse = -scores.mean()
  • In scikit-learn every score is "higher is better", so error metrics come back negated: flip the sign to read the RMSE
Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation, back to concrete

  • The 86 test mixes (195 rows) were locked away first; the models see the other 835 rows
  • A row is one mix crushed at one age, and 182 mixes were crushed at several ages
  • So rows are not independent experiments: two rows can be the same mix
Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation, the four families on concrete

5-fold cross-validation on the 835 rows, rows assigned to folds at random:

Model RMSE (MPa) MAE (MPa)
Baseline: predict the mean 16.8 13.6 −0.01
Linear 10.6 8.46 0.60
Linear, with physics features 7.25 5.60 0.81
Decision tree (no depth limit) 6.92 4.52 0.83
Neural network (16 tanh) 5.96 4.20 0.87
Gaussian process 5.81 3.96 0.88

Physics features: log(age), and the water/cement ratio

Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation, grouped rows

Grouped cross-validation (GroupKFold): all the rows of a group (a mix, a batch, a run) are either all in training or all in validation.

  • The engineer's question: how strong will a mix nobody has made yet be? Only grouped folds ask it
Systems and Toolchains for AI Engineers
06-763 / L9

Cross-validation, KFold against GroupKFold

  • Random rows: the tree beats the line with physics features (6.92 against 7.25 MPa)
  • Whole mixes: the tree loses by 2 MPa (9.42 against 7.43); the line is within 0.3 MPa of the GP
Systems and Toolchains for AI Engineers
06-763 / L9

Model capacity, overfitting and learning curves

Systems and Toolchains for AI Engineers
06-763 / L9

Model capacity

Capacity: the range of functions a model can represent. More capacity fits more complicated relationships, and more of the noise.

  • The knobs: polynomial degree, tree depth, hidden units, the penalty , a kernel's length scale
  • Overfitting: training error much lower than validation error. Underfitting: both high, and close together

Bias-variance trade-off: more capacity lowers bias and raises variance, so validation error is lowest in between.

Hastie, Tibshirani and Friedman, section 7.3, eq. 7.9

Systems and Toolchains for AI Engineers
06-763 / L9

Model capacity, validation curves

Validation curve: training and validation error against one capacity knob, with the data fixed.

Systems and Toolchains for AI Engineers
06-763 / L9

Model capacity, learning curves

Learning curve: training and validation error against training-set size, model fixed.

  • Gap is variance: the line's gap closes to 0.3 MPa, so it no longer overfits
  • Level is bias plus noise: the line's validation ends at 7.4 MPa, more capacity reaches 6.1, so it underfits
  • The tree: a gap of 8.5 MPa that more samples are not closing, so it overfits
The curves show Diagnosis What helps Here
Gap closed, error still high Underfitting (high bias) More capacity, better features The line (left)
Gap closed, error low A good fit Stop, and test once
Big gap: training low, validation much higher Overfitting (high variance) Less capacity, regularization The tree (right)
Validation still falling at the right edge Limited by data More training samples
Systems and Toolchains for AI Engineers
06-763 / L9

Limitations

Systems and Toolchains for AI Engineers
06-763 / L9

Limitations, outside the data

All four refitted on the 21 points. At 300 °C, NIST: 517.7 MPa. Polynomial 192, tree flat at 100.7, network saturates at 148, GP 162 ± 232: the band still misses.

Systems and Toolchains for AI Engineers
06-763 / L9

But wait, we didn't discuss the hyperparameters?

Chosen by hand today: 16 hidden units, a tree with no depth limit, the ridge of the ARX, the kernel's form, and the depth-2 water tree

  • An optimization problem with a training problem inside it
  • Every evaluation of the outer objective trains a model
  • The best of many validation scores is optimistic

How do you search over without fooling yourself?

Systems and Toolchains for AI Engineers
06-763 / L9

Worked example

l09-regression.ipynb

The workflow on concrete, one step per cell: lock the test mixes, fit four families, compare KFold with GroupKFold, test once. Then NARX on Lecture 8's table.

Run it after class, top to bottom. The first run downloads the data; the cross-validation cell takes a minute or two.

Open the worked example

Systems and Toolchains for AI Engineers
06-763 / L9

Recap

Four questions to ask of any model, today's or your own:

What does it know?

Only its data.

The polynomial scored R2 = 0.9999975 on the water data, and gave 223 MPa at 300 °C, where NIST gives 517.7.

What did training solve?

An optimization problem, and the family picks it.

A line has one minimum and a network many; a GP tunes a few bounded hyperparameters.

What did the score measure?

The question your split asked.

The tree beat the line on random folds and lost to it by 2 MPa on grouped folds.

What is holding it back?

A gap is variance, a high level is bias.

The tree kept a gap of 8.5 MPa. The line closed its gap but ended 1.4 MPa above a model with more capacity.

Start simple, and put what you know into the features: a line with two physics features came within 0.3 MPa of a Gaussian process on grouped folds.

Systems and Toolchains for AI Engineers
06-763 / L9

This week

Practice module for this session, for participation credit
Worked example l09-regression.ipynb, to run after class

Systems and Toolchains for AI Engineers

110 minutes: lecture 90, questions 20. The notebook is a worked example students run after class; it is not shown live. Plan, by slide: opening and the examples (2-7) 9, types (9-12) 5, workflow (14-17) 6, training as optimization (19-23) 9, regression families (25-52) 36, cross-validation (54-59) 7, capacity (61-63) 5, limitations and the close (65-69) 5. About 82 minutes at a steady pace. If it runs long, the first to go, in order: the optimizer paths (23, one sentence), what each training solves (48), the kernel and prediction slide (43, the notes carry it), the NARX forecast figure (52).

A pause for the comic. "Just stir the pile until they start looking right" is the failure this session is about. ML: building models that learn patterns from data; instead of explicitly programming rules of physics/nature, we train algorithms to generalize from examples.

"Constant density" is the physics: a fixed mass of liquid water in a sealed, rigid container, so its density stays at 1000 kg/m3. It cannot expand when heated, so the pressure climbs steeply, from about 0.4 MPa near 0 C to about 100 MPa (roughly 1000 atm) at 100 C. First of today's examples; it comes back for feature engineering, regularization, the first tree, and at the end for what every family does outside its data.

The paper's system: the surfactant cetylpyridinium chloride (also the antiseptic in some over-the-counter mouthwashes) with the salt sodium salicylate. The long worms are wormlike micelles, and they tangle like polymer chains. Zero-shear viscosity is the plateau the viscosity reaches as the shear rate goes to zero: the thickness of the liquid sitting still. Why the peak: the salt screens the charges on the surfactant heads, so small spherical micelles grow into long worms and the viscosity climbs; past the peak it falls, commonly attributed to the worms branching or shortening (which of the two is still debated, Ziserman et al. 2009). The Kitchin group page labels the axis salt concentration and gives no units. The 16 points follow the shape of the paper's curve (sharp peak, dip, smaller second peak, fall) on a compressed scale: about 400-fold here, about five decades in the paper's own curve (replotted in Berret 2004, at 100 mmol/L of surfactant).

Sand and gravel are the file's fine and coarse aggregate. Slag (a byproduct of iron making) and fly ash (from coal power plants) replace part of the cement; superplasticizer is an additive that lets the fresh mix flow with less water. In standard practice a test result is the average of at least two cylinders crushed at the same age (NRMCA CIP 35); the UCI files do not say how many cylinders each row averages. Crushing destroys a cylinder, so a line is several cylinders of one mix, one per age. Designs specify strength at 28 days, which is why 425 of the 1,030 rows are 28-day tests. Only 428 distinct mixes, and 182 of them were crushed at several ages, which matters for the folds in the cross-validation section. High-performance concrete: concrete meeting performance and uniformity requirements that conventional ingredients and practice cannot always achieve (ACI).

Supervised: labels (target outputs). Unsupervised: no labels, find hidden structure (clustering, dimensionality reduction). Reinforcement: actions, state updates, feedback. Today is all supervised.

"ML is a bit jargonized. Let's clarify most of the terms used before we move on."

Four cards: what is one row, what goes in, what comes out, which task.

1 Feature engineering: select or transform the inputs (polynomial features of T, or the lagged columns of a NARX table: Lecture 7's regressors are features). Anything fitted to data, such as a scaler, is fitted after the split, on the training rows only. 2 Data splitting: training set to fit, test set to evaluate on unseen data. 3 Model selection: choose the algorithms. 4 Model validation: prevent overfitting, with metrics. Then the test set, once.

scikit-learn testimonials (https://scikit-learn.org/stable/testimonials/testimonials.html): is it used in real applications? Yes. The next section's question: what does each .fit solve?

The labels and their arrows come in one at a time.

The labels and their arrows come in one at a time. Same shape as the previous slide, now with the model's parameters as the decision variables and the data as constants. Regularization adds a penalty term to this objective, later in the regression section.

Four terms, one at a time. This one update covers every method on the next slide; they differ only in where g comes from and what H is.

One update for all of them: step = -eta H g. Gradient descent: H = I and the full gradient. SGD: H = I and the gradient of a random mini-batch, a noisy estimate of the full one; cheap steps, and the step length has to shrink for the iterates to settle. Adam: SGD plus two running averages, of g (momentum) and of g squared (a step size per parameter), bias-corrected because both start at zero; defaults 0.001, 0.9, 0.999. L-BFGS: quasi-Newton; keeps the last m pairs (s = the step, y = the change in gradient) instead of a Hessian, builds H g from them, line search; needs an exact full gradient, so small and medium data. scikit-learn's MLPRegressor docs: adam "works pretty well on relatively large datasets (with thousands of training samples or more)"; "for small datasets, however, 'lbfgs' can converge faster and perform better". The GP: L-BFGS-B, with bounds on every hyperparameter.

One clock for all four, the iteration number, on a log scale, so it speeds up as it runs: L-BFGS's six steps fill the first three seconds and the last 1,600 iterations take about three. A dot drops out when its method reaches the minimum, and its legend entry turns bold. SGD is drawn at every step to 400, then every 4th. A least squares problem, so the loss is a quadratic bowl, but a long narrow one: intercept and slope trade off (condition number 37.7). L-BFGS learns the shape of the valley in its first steps. SGD (4 rows per step, fixed step) reaches the floor fast and keeps bouncing: every mini-batch points somewhere slightly different. Adam runs on the full gradient here, so only its update differs from gradient descent; its momentum overshoots (the loop), then it settles.

"Merely a polynomial model, so you should not use it for extrapolation." Fitted with a library that implies ML, but no physics in it. At -50 C: 45.6 MPa. At 300 C: 223 against 517.7. Within the data range, a reasonable estimation.

Large alpha: a simpler model, and a risk of underfitting. Small alpha: plain least squares.

Small alpha: the curve chases the noise near the edges, training RMSE low, test RMSE higher. Large alpha: the curve flattens, both RMSEs rise. Lasso: the nonzero count falls as alpha grows. The right panel is a validation curve.

Linear regression first, because it is the one they know; this is why the other three exist. "Every possible problem" means every input-output relationship, weighted equally, and "new inputs" means points outside the training set (Wolpert's off-training-set error). Strictly, the 1996 theorem is proved for losses like zero-one, where a guess is right or wrong. For squared error the companion paper finds an edge, but all of it comes from where a method's guesses sit in the output range; on average, the way a method uses its data adds nothing. The surprise in Wolpert's abstract: it holds even for cross-validation against "anti-cross-validation" (pick the model with the largest validation error). So the second sentence is our reading and goes beyond the theorem: trusting validation is itself an assumption, that real problems have structure (smoothness, boxes, the right features) and that held-out points resemble the ones you will predict.

"Something interesting is happening here." Each split is the boundary that gives the minimum MSE. Depth 2 means four leaves, four values, a staircase. Test R2 0.911. Trees capture nonlinearity but overfit if the depth is too large, so max_depth is crucial. Finding the optimal tree is NP-complete (Hyafil and Rivest 1976), which is why it is grown greedily: no gradient, no starting point. Laird group: linear model decision trees inside optimization (OMLT).

One hidden unit at a time lights up and its term appears beside it, at the same height, in its color; then the output bias; then the sum. The previous slide's equation, one row per unit. The input weight w0k and bias b0k sit inside the tanh, the output weight w1k outside it. The deep version, several hidden layers, is on the hyperparameters slide.

Each colored curve is one unit from the previous slide: w0k and b0k set where the curve bends and how sharply, w1k sets how tall it is and its sign. The fitted terms are large and nearly cancel (offsets near -37, -23 and +64), so each is drawn shifted to start at zero; the offsets and b1 fold into one constant. The sum is the fit.

Play or the slider moves x. A unit lights up when it is on. Each kink is one unit switching on or off, so a ReLU network is piecewise linear. Unit 3 is on over the whole range, so it only adds a straight line.

Gradients by backpropagation. The worked example silences a ConvergenceWarning from the concrete network: the optimizer hit its 5,000-iteration limit before declaring convergence.

With x2 in the millions, all 20 tanh units saturate at plus or minus 1 on every training row, the gradients vanish and the network predicts the training mean (0.128) everywhere. Standardizing fixes it: zero mean, unit variance, fit on training rows only.

The picture from my F25 GP slides, built in four steps. It starts on a distribution of numbers, N(0, 1). First, fifteen draws land one at a time, each one a number; the newest is red. One, at -2.33, falls outside 2 std, where about 5% of draws should (4.55%). Second, the same idea one level up: the arrow, then a distribution of functions, its mean (dashed) and a band of 2 std. On the left the same range turns gray: at every x this GP's value is N(0, 1), the bell on the left, so the band is that gray range repeated at every x. Third, five draws, each a whole function, traced one after another. The gold one leaves the band at both ends: the band covers about 95% of the values at each x, so a whole curve can still cross. Fourth, the definition. How fast the curves wiggle is the kernel's length scale (0.3 here), two slides on. Formally (Rasmussen and Williams, Definition 2.1): any finite set of the function's values is jointly Gaussian. Parametric (a network): fix theta and fit f(x; theta). Nonparametric (a GP): a distribution over functions, and every prediction comes with an uncertainty. Going back a step undoes it; going forward again replays it.

Play or the slider adds the points one at a time. At 0 it is the prior, every function the kernel allows: mean 0 and a band of plus or minus 4.12 (2 std) everywhere. Each point is a noisy sample of f(u) = sin(u) + log(u) - exp(-0.1 u^2); the newest has a red ring. The band pinches at each point and stays wide in the gaps: average half-width 4.12, then 2.22 after 2 points, 0.80 after 5 and 0.16 after 20. The RMSE of the mean falls unevenly: near 0.45 from 7 to 9 points, then 0.083 at 10, when point 10 (u = 0.76) fills the gap at the left edge. The kernel (signal std 2.06, length scale 2.14, noise std 0.112) was fitted once on all 20 points and held fixed, so only the data change. Bayesian inference conditions on the data: the mean is a weighted average of the observed outputs, and the variance is what the data have not pinned down.

The squared exponential (RBF) kernel. The mean is a weighted average of the known outputs, with weights set by similarity; the variance starts at the prior variance and drops by what the data explain, so it grows far from the data. Other kernels: Matern (rougher), periodic, sums and products.

Strengths: probabilistic predictions, interpretable kernels, automatic Occam's razor. Weaknesses: O(N^3), kernel choice matters, needs scaling.

Checked in the scikit-learn source (1.6 and 1.9): GaussianProcessRegressor calls scipy.optimize.minimize(method="L-BFGS-B", bounds=kernel.bounds). Theta and the bounds are log-transformed, default (1e-5, 1e5) for the length scale, the constant (signal variance) and the white-noise level. alpha goes on the diagonal and is not optimized. Restarts are drawn log-uniformly inside the bounds. MLPRegressor(solver="lbfgs") calls the same routine without bounds; adam and sgd apply none. GPflow: L-BFGS-B on softplus-transformed parameters. GPyTorch: Adam on raw parameters with a softplus positivity constraint.

The labels and their arrows come in one at a time. The marginal likelihood averages over every function the GP prior allows, so a flexible kernel spreads its probability over many possible datasets and gives little to the one measured: that is the log-determinant term. My F25 slide: "GP estimation balances data fit and complexity of the predicted model: Occam's razor is automatic".

Short l: the mean spikes through every point and falls back to the average between them; complexity penalty large. The function may change between neighboring points. The sum reaches its minimum near l = 0.31 (the fitted value on all 16 points; the 0.235 two slides back is in standardized units, on 12 points). Long l: smooth over the whole range; the curve cannot reach the peak and the misfit explodes.

Four bullets, one at a time. F is any residual known to be zero: a balance, a rate law, an ODE evaluated at collocation points z_j. One sentence is enough.

Same fit/predict, on a lag table instead of a table of experiments. The GP-NARX got only 1,000 rows because its cost grows as N cubed: all 144,300 training rows would need a 167 GB kernel matrix (the notes have a section on where scikit-learn stops). ARX on the same 1,000 rows: 4.76. Lecture 7's deviation variables are the same idea as scaling the target.

Random folds: the 28-day row of a mix trains and the 56-day row of the same mix validates, so the model is asked about a mix it has already seen. Lecture 8's split by run, without the time axis.

ESL eq. 7.9: Err(x0) = noise + Bias^2 + Var. Noise: no model removes it. Bias: the average model's distance from the truth. Variance: how much the model moves when the training set changes. A straight line on concrete is the high-bias end; a tree with no depth limit is the high-variance end.

Training error falls to 0.95 and never to zero: nine settings (same mix, same age) were crushed more than once with different results, one from 22.9 to 55.9 MPa at 7 days. Those replicates are the noise term. GroupKFold validation stops improving at depth 9 (9.10) and wobbles 9.2 to 9.6. "Looking only at the train score can be misleading!" Random folds keep rewarding depth because they reward memorizing.

Each panel shows two things. The gap between the curves is the variance: the line's closes to 0.3 MPa as the training set grows, so it no longer overfits, and more samples can buy at most 0.3. The level where the curves meet is bias plus noise. "High" needs a reference: gradient-boosted trees on the same features and grouped folds reach 6.1 MPa, better than the line on all five folds. So about 1.4 MPa of the line's 7.4 on validation is bias that more capacity removes; the rest includes the replicate noise from the last slide. The line underfits. The same closed gap at a level nothing beats would be a good fit: stop and test once. The tree: 0.95 on its training samples, 9.4 on validation, a gap of 8.5: overfitting. Its validation curve is flat from about 400 samples on, so more samples of the same kind do not close the gap; hence no example here for the "limited by data" row. The curves average ten random orderings of the training rows. learning_curve does not shuffle by default, and in file order the tree's curve showed a false late drop.

A GP's uncertainty describes distance from the data under the kernel's smoothness assumption. It says nothing about physics the data never showed it. The notes' four-families table has a row for this: follows its features, flat, saturates, back to the mean. 192 here, not the 223 of the opening slide: this polynomial is fitted on all 21 points, that one on 16.

Not run in class; the 20 minutes after the deck are for questions. The notebook: the concrete workflow and the NARX forecasts from today's slides, one step per cell, on the real data. The first run downloads the concrete file (125 kB) and the fault-free TEP file (25 MB).

Each card's picture is the slide where the class saw the answer. The four questions hold for any model the students fit this semester, so the deck ends on them instead of a list. The habits under them: lock the test set and touch it once, scale the inputs and the target, set random_state.