A model overfits when its training error is much lower than its validation error. It has learned the noise and the particular rows it was given. It underfits when both errors are high and close together, because it is too simple for the relationship.
The bias-variance trade-off: more capacity lowers bias and raises variance, so the validation error is lowest somewhere in between.
A validation curve plots training and validation error against one capacity knob, with the data held fixed.

A learning curve plots training and validation error against training-set size, model fixed.

| The curves show | Diagnosis | What helps | Here |
|---|---|---|---|
| Gap closed, error still high | Underfitting (high bias) | More capacity, better features | The line (left) |
| Gap closed, error low | A good fit | Stop, and test once | |
| Big gap: training low, validation much higher | Overfitting (high variance) | Less capacity, regularization | The tree (right) |
| Validation still falling at the right edge | Limited by data | More training samples |
Classification predicts a category, usually by first predicting a probability for each class.
The Tennessee Eastman process (TEP), the simulated plant from Lecture 5 on. The question: is the plant faulty right now?
One sample: one 3-minute snapshot of the plant
Features: the 52 channels at that moment
Target: normal (0) or faulty (1)
Split by whole run, as in Lectures 8 and 9
Fault 4: the cooling water entering the reactor gets warmer, so the controller opens the cooling water valve, xmv_10, further. P&ID: Lyu et al. (2026)
Logistic regression turns a weighted sum of the inputs into a probability between 0 and 1.
xmv_10) times a weight 
Below 43.4% open: normal. Above: fault 4. On 51,000 test samples, this one cut is wrong once.
Fault 14: the same valve sticks. The controller pushes, the valve jumps too far, the controller pushes back: big swings around the same average.
P&ID: Lyu et al. (2026)

The fault samples fall on both sides of the normal band. One straight cut cannot separate that: logistic regression catches 0%. A tree's two cuts flag both sides and catch 88%.

Step 3000
Minimizes the log loss, also called cross-entropy: small when the true class gets a high probability.

Depth 3
Splits to lower the Gini impurity of each box: zero when a box holds one class.

Trained (172 iterations)
A softmax turns its scores into probabilities; training minimizes the cross-entropy.

150 points
A smooth random function
Accuracy: the fraction of predictions that are correct.
A confusion matrix counts, for every sample, what really happened against what the model said. Call a fault a positive and a normal sample a negative.
The numbers: a neural network on all 52 channels, trained on normal runs and nine faults, tested on the 59,000 test samples (any of the nine faults counts as a fault), threshold 0.5.
Precision: the share of the model's alarms that were real faults. Recall: the share of the real faults that the model caught.
| Classifier, on the 59,000 test samples | Accuracy | Precision | Recall |
|---|---|---|---|
| Always normal | 0.854 | 0 | 0 |
| Logistic regression, 52 channels | 0.970 | 0.993 | 0.803 |
| Decision tree, 52 channels | 0.977 | 0.988 | 0.855 |
| Neural network, 52 channels | 0.994 | 0.998 | 0.961 |
The slider moves the threshold on the network's probability of fault, to show what each choice costs. In scikit-learn, model.predict(X) gives the class (threshold 0.5); model.predict_proba(X) gives the probability, so you pick the threshold.
Threshold ↓: more alarms, recall ↑, precision ↓. Threshold ↑: fewer alarms, precision ↑, recall ↓.

A classifier only catches new faults that look like the ones it was shown.
Lecture 9 picked the tree's depth by eye. Its RMSE (root mean squared error) over grouped folds: 9.42 MPa with no depth limit, 9.10 MPa at depth 9.
Optuna: an open source Python library that searches hyperparameters for you.
You write an objective function; Optuna proposes the trials and keeps the best. MLflow, from Lectures 1 and 2, records runs. A search uses both:
Every trial stays on record, and the best model is registered.
Two hyperparameters, nine trials for each search. Only the x-axis hyperparameter matters in this example.
The bump (made up for the picture): the validation score against the x-axis hyperparameter. Its peak is the best setting.
Takeaway: same 9 trials, but random search tries 9 values of what matters and grid only 3, so random gets closer to the best.
Bergstra and Bengio (2012), JMLR 13TPE (Tree-structured Parzen Estimator): Optuna's default rule for picking the next trial.

Bergstra et al. (2011), where TPE comes from
Each dot: one trial's validation RMSE. Each line: the best RMSE so far, so it only goes down.
Green dotted lines: Lecture 9's two trees, 9.42 and 9.10 MPa. TPE's best, 8.90 MPa, beats both.
Akiba et al. (2019), the Optuna paperA search trains a model once per trial. MLflow keeps every one of them, in one place:
Each child run is one trial: a new max_depth and min_samples_leaf, and the validation RMSE they give. The data, the folds and the model stay the same. The registry holds the model to use, with a name and a version anyone can load.
def objective(trial):
params = dict(
max_depth=trial.suggest_int('max_depth', 2, 20),
min_samples_leaf=trial.suggest_int('min_samples_leaf', 1, 50),
)
model = DecisionTreeRegressor(
random_state=SEED,
**params,
)
rmse = -cross_val_score(
model,
X_tr,
y_tr,
groups=groups_tr,
cv=GroupKFold(5),
scoring='neg_root_mean_squared_error',
).mean()
with mlflow.start_run(nested=True):
mlflow.log_params(params)
mlflow.log_metric('val_rmse', rmse)
return rmse
def objective(trial): Optuna calls it once per trial. It returns the score to minimize.trial.suggest_int(name, low, high): Optuna picks an integer in that range for this trial.random_state=SEED fixes its randomness, **params passes this trial's settings.cross_val_score, as in Lecture 9: cv=GroupKFold(5) makes five folds, groups= keeps each mix in one fold, scoring= asks for the RMSE (negative, hence the minus sign).mlflow.start_run(nested=True): a child run inside the parent run, with this trial's settings and score.return rmse: the number Optuna minimizes.parent = mlflow.start_run(run_name='concrete-tree-search')
study = optuna.create_study(
direction='minimize',
sampler=optuna.samplers.TPESampler(seed=SEED),
)
study.optimize(
objective,
n_trials=20,
)
winner = DecisionTreeRegressor(
random_state=SEED,
**study.best_params,
).fit(X_tr, y_tr)
mlflow.sklearn.log_model(
winner,
name='model',
registered_model_name='concrete-tree',
skops_trusted_types=['sklearn.tree._tree.Tree'],
)
mlflow.end_run()
mlflow.start_run(run_name=...): opens the parent run. It stays open for the whole search.create_study: direction='minimize', since a lower RMSE is better. sampler= picks TPE, seeded so the search repeats.study.optimize: runs 20 trials, so 20 child runs.study.best_params, refit on all 835 training rows.log_model: saves and registers the tree as concrete-tree. skops_trusted_types= marks its type as safe to load; MLflow requires it.mlflow.end_run(): closes the parent run.concrete-tree, version 1Report the test number, 7.41 MPa. The test mixes never took part in the search.
It is lower than 8.90 here: 195 rows from 86 mixes is a small test set, and its score depends on which mixes landed in it.
l10-classification.ipynb: the 52-channel classifier, its confusion matrix, three thresholds, and its recall on faults it never saw.
l10-tracking-search.ipynb: a 20-trial Optuna search on the concrete strength dataset, one MLflow child run per trial, the winner registered and tested once.
Four questions this session added to Lecture 9's four:

A threshold on a probability.
Fault 4's boundary sits at 43.4% open; fault 14 needs more than one straight cut.

Its confusion matrix, not accuracy alone.
The baseline scores 85.4% accuracy and 0 recall; precision and recall expose it.

New faults only if they look like the ones it was shown.
Two faults it never saw: fault 18 was caught 92% of the time, fault 19 only 0.1%.

Track every trial, test the winner once.
TPE reached 8.90 MPa at trial 11, random search 9.00 at trial 38. The test set gave 7.41.
Report the confusion matrix with its threshold, track every trial, and test once.
Practice module for this session, for participation credit
Demos l10-classification.ipynb and l10-tracking-search.ipynb, to rerun after class
Last time 10, classification 25, measuring a classifier 20, tracking and search 20, demos 10. That is 85 minutes of deck, with 25 left after it.
The plant carries classification and its measures. The concrete strength dataset comes back for the search, with Lecture 9's own decision tree.
Ten minutes before anything new. Same four questions, then capacity again, slower, because I rushed the end of last time.
Same four cards from the end of last time: what a model knows, what training solved, what the score measured, what was holding it back. These apply to the classifier we build today too, so I want them fresh before we start.
Same five knobs as last time, named one at a time: degree, depth, hidden units, alpha, length scale.
Overfit and underfit again, plus the bias-variance line underneath them, since I went past this fast last time. A straight line on the concrete strength dataset is the high-bias end, an unlimited tree the high-variance end.
One figure, one story: the black training curve keeps falling, the red validation curve flattens past depth nine. Looking only at the training score is misleading.
Gap is variance, level is bias. The line's gap closed but its level is still a bit high, so it underfits. The tree's gap never closed. Same four rows as last time, reread slower.
Same fit and predict as Lecture 9. What changes is the output: a probability first, then a threshold turns it into a class.
One channel and one fault first: the valve sits near 41% open in normal runs and near 45% under fault 4. All 52 channels come in later, for the classifier we measure.
Same weighted sum as linear regression. The fraction and the threshold are the new parts. 0.5 is only a starting choice; the threshold slider moves it later.
The boundary is where the weighted sum is zero, so x = -b/w. Normal runs average 41.1% open, fault 4 runs 44.9%, and they barely overlap. The one mistake is a single false alarm.
Stiction: friction holds the valve until the push breaks it free, then it overshoots, and the loop settles into an oscillation. The controller still holds the reactor temperature on average, so the mean stays near 41% open. xmv_10 is the command to the valve, so these swings are the controller fighting the stuck valve.
The normal samples sit in a narrow band, the fault 14 samples spread from about 29% to 54%. A tree of depth 8 catches 99.7%.
Same fit and predict for all four. Each card replays its model learning: logistic regression by gradient descent from zero, the tree one depth at a time, the network iteration by iteration, the Gaussian process as the points arrive. A line, boxes, a smooth curve, smooth probabilities.
A plant runs normally most of the time, so answering "normal" is right most of the time. Accuracy cannot tell this detector from a useful one.
8,300 faults caught, 340 missed, and only 13 false alarms among 50,360 normal samples. Top left and bottom right are right; the other two boxes are the two kinds of mistake.
Precision: when it raises an alarm, can I trust it? Recall: of the real faults, how many did it catch? Both numbers come from the same four boxes.
The baseline is the "always normal" detector from the accuracy slide. Logistic regression misses one fault sample in five; the network misses one in twenty-five.
At 0.01, 8,438 of the 8,640 faults are caught, with 3,664 false alarms. At 0.99 there are no false alarms, and 7,917 are caught. Recall moves little and precision a lot, because this classifier already separates the two classes well.
The blue bar is the nine faults it trained on, 0.961. Nothing in the classifier says in advance which new faults it will catch.
Choosing hyperparameters is an optimization problem with training inside it. The outer problem is what a search solves.
Two jobs: Optuna decides what to try next, MLflow writes down what happened.
Nine trials either way. Grid repeats each value of the important hyperparameter three times, so it tries only 3 values of it. Random almost never repeats a value, so its nine trials cover that axis better.
The first 20 trials of the search on the concrete strength dataset. Good means the best 10%: 2 of 20, both with leaf size 9. The bad ones spread from 2 to 47. Trials 21 to 26 use leaf sizes 8 to 13. Optuna's own curves are smoother than these.
Trials 1 to 10 are the same for both: TPE starts with 10 random trials. From trial 11, TPE tries leaf sizes of 2 to 20, where the good trials were; random keeps drawing 21 to 48. TPE reaches 8.90 at trial 11, random gets to 9.00 only at trial 38.
One run per trial, as in Lectures 1 and 2, now grouped under a parent. Every child run stays on record. The registry is where the chosen model goes.
The same objective as the notebook, one piece at a time. Everything inside it is Lecture 9's grouped cross-validation; the two new parts are suggest_int and the nested run.
The parent run is opened before the study and closed after the registration, so the twenty child runs, the best settings and the registered model all sit under it.
Same rule as Lecture 9, now for the winner of a search: the test set is touched once, after the search is over.
About five minutes each. The first run of the classification notebook downloads 45 MB of plant data.
Each card's picture is the slide where the room saw the answer. The closing line is the three habits under the four cards.