06-763 / L10

Lecture 10: The machine learning workflow III, classification, tracking and search

Week 5, Machine learning and deep learning

Systems and Toolchains for AI Engineers

Systems and Toolchains for AI Engineers
06-763 / L10

What today is about

  1. Last time, slower: model capacity, validation curves, learning curves
  2. Classification: predicting a category, on the plant
  3. Measuring a classifier: the confusion matrix, precision and recall, the threshold
  4. Tracking and search: Optuna runs the search, MLflow records every trial
Systems and Toolchains for AI Engineers
06-763 / L10

Last time

Systems and Toolchains for AI Engineers
06-763 / L10

Last time, four questions to ask of any model

Four questions to ask of any model, last week's or your own:

What does it know?

Only its data.

The polynomial scored R2 = 0.9999975 on the water data, and gave 223 MPa at 300 °C, where NIST gives 517.7.

What did training solve?

An optimization problem, and the family picks it.

A line has one minimum and a network many; a GP tunes a few bounded hyperparameters.

What did the score measure?

The question your split asked.

The tree beat the line on random folds and lost to it by 2 MPa on grouped folds.

What is holding it back?

A gap is variance, a high level is bias.

The tree kept a gap of 8.5 MPa. The line closed its gap but ended 1.4 MPa above a model with more capacity.

Start simple, and put what you know into the features: a line with two physics features came within 0.3 MPa of a Gaussian process on grouped folds.

Systems and Toolchains for AI Engineers
06-763 / L10

Model capacity

Capacity is the range of functions a model can represent. More capacity fits more complicated relationships, and more of the noise.

Every family from last time turns a different knob:

  • Polynomial degree, for the line
  • Tree depth
  • The number of hidden units, for the network
  • The ridge or lasso penalty
  • A kernel's length scale, for the Gaussian process
Systems and Toolchains for AI Engineers
06-763 / L10

Model capacity, overfitting and underfitting

A model overfits when its training error is much lower than its validation error. It has learned the noise and the particular rows it was given. It underfits when both errors are high and close together, because it is too simple for the relationship.

The bias-variance trade-off: more capacity lowers bias and raises variance, so the validation error is lowest somewhere in between.

Hastie, Tibshirani and Friedman, section 7.3, eq. 7.9

Systems and Toolchains for AI Engineers
06-763 / L10

Model capacity, validation curves

A validation curve plots training and validation error against one capacity knob, with the data held fixed.

  • Training error keeps falling as a tree gets deeper, and validation error stops improving past depth 9
Systems and Toolchains for AI Engineers
06-763 / L10

Model capacity, learning curves

A learning curve plots training and validation error against training-set size, model fixed.

  • The gap between the curves, at the right edge, is variance
  • The level where they meet is bias, plus noise
  • A closing gap is good news: more data has already done its job
  • The level tells underfitting from a good fit; the gap only tells you whether more data would help
The curves show Diagnosis What helps Here
Gap closed, error still high Underfitting (high bias) More capacity, better features The line (left)
Gap closed, error low A good fit Stop, and test once
Big gap: training low, validation much higher Overfitting (high variance) Less capacity, regularization The tree (right)
Validation still falling at the right edge Limited by data More training samples
Systems and Toolchains for AI Engineers
06-763 / L10

Classification

Systems and Toolchains for AI Engineers
06-763 / L10

Classification, what it is

Classification predicts a category, usually by first predicting a probability for each class.

  • Regression (Lecture 9) predicts a number: a strength, a pressure
  • Classification predicts a category: normal or faulty, pass or fail
  • Most classifiers work in two steps: a probability, then a threshold
52 plant channels
at one moment
→
The model
→
A probability,
P(fault)
→
The threshold:
is P(fault) ≥ 0.5?
→
Fault or normal
Systems and Toolchains for AI Engineers
06-763 / L10

Classification, the plant

The Tennessee Eastman process (TEP), the simulated plant from Lecture 5 on. The question: is the plant faulty right now?

One sample: one 3-minute snapshot of the plant

Features: the 52 channels at that moment

Target: normal (0) or faulty (1)

Split by whole run, as in Lectures 8 and 9

Fault 4: the cooling water entering the reactor gets warmer, so the controller opens the cooling water valve, xmv_10, further. P&ID: Lyu et al. (2026)

Systems and Toolchains for AI Engineers
06-763 / L10

Classification, logistic regression

Logistic regression turns a weighted sum of the inputs into a probability between 0 and 1.

p = 1 1 + e−(wx + b) p ≥ 0.5 → faultp < 0.5 → normal
  • Weighted sum : the valve opening (xmv_10) times a weight plus an offset
  • The fraction squashes any number into a probability between 0 and 1
  • Threshold at 0.5: call it a fault when is at least 0.5, otherwise normal
Systems and Toolchains for AI Engineers
06-763 / L10

Classification, one cut on the valve

  • Blue curve: the fitted at each valve opening
  • Ticks: training samples, normal at the bottom, fault 4 at the top
  • Training found and
  • Boundary: where open

Below 43.4% open: normal. Above: fault 4. On 51,000 test samples, this one cut is wrong once.

Systems and Toolchains for AI Engineers
06-763 / L10

Classification, fault 14

Fault 14: the same valve sticks. The controller pushes, the valve jumps too far, the controller pushes back: big swings around the same average.

P&ID: Lyu et al. (2026)

Systems and Toolchains for AI Engineers
06-763 / L10

Classification, fault 14, one cut or two

The fault samples fall on both sides of the normal band. One straight cut cannot separate that: logistic regression catches 0%. A tree's two cuts flag both sides and catch 88%.

Systems and Toolchains for AI Engineers
06-763 / L10

Classification, four model families

Step 3000

Logistic regression

Minimizes the log loss, also called cross-entropy: small when the true class gets a high probability.

Depth 3

Decision tree

Splits to lower the Gini impurity of each box: zero when a box holds one class.

Trained (172 iterations)

Neural network

A softmax turns its scores into probabilities; training minimizes the cross-entropy.

150 points

Gaussian process

A smooth random function goes through a link . It also gives an uncertainty.

Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier

Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, accuracy

Accuracy: the fraction of predictions that are correct.

  • Imagine a "detector" that answers normal for every sample, whatever the data. On the 59,000 test samples:
85.4% of the samples are normal: the detector is right on all of them
14.6% faulty:
all missed
  • Accuracy 85.4%, yet it never catches a single fault
  • So we need metrics that tell the two mistakes apart: missed faults and false alarms
Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, the confusion matrix

A confusion matrix counts, for every sample, what really happened against what the model said. Call a fault a positive and a normal sample a negative.

The numbers: a neural network on all 52 channels, trained on normal runs and nine faults, tested on the 59,000 test samples (any of the nine faults counts as a fault), threshold 0.5.

What the model said: fault What the model said: normal What really happened: fault What really happened: normal True positive (TP) fault caught 8,300 False negative (FN) fault missed 340 False positive (FP) false alarm 13 True negative (TN) left alone 50,347
  • Read across a row: what really happened
  • Read down a column: what the model said
  • The diagonal, TP and TN: the model was right
  • Off the diagonal, FN and FP: the two kinds of mistake
Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, precision and recall

Precision: the share of the model's alarms that were real faults. Recall: the share of the real faults that the model caught.

  • Why two numbers: an alarm you cannot trust (low precision) and a fault you miss (low recall) cost different things
  • The network: of its 8,313 alarms, 8,300 were real faults, so precision is 0.998
  • Of the 8,640 real faults, it caught 8,300, so recall is 0.961
Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, four classifiers

Classifier, on the 59,000 test samples Accuracy Precision Recall
Always normal 0.854 0 0
Logistic regression, 52 channels 0.970 0.993 0.803
Decision tree, 52 channels 0.977 0.988 0.855
Neural network, 52 channels 0.994 0.998 0.961
  • Always normal never predicts a fault, so TP = 0 and recall is 0
  • It raises no alarm at all, so precision is 0/0, which scikit-learn reports as 0
  • Accuracy barely separates the last three rows. Recall does.
Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, the threshold

The slider moves the threshold on the network's probability of fault, to show what each choice costs. In scikit-learn, model.predict(X) gives the class (threshold 0.5); model.predict_proba(X) gives the probability, so you pick the threshold.

Threshold on P(fault) t = 0.50

Threshold ↓: more alarms, recall ↑, precision ↓. Threshold ↑: fewer alarms, precision ↑, recall ↓.

Systems and Toolchains for AI Engineers
06-763 / L10

Measuring a classifier, faults it never saw

  • The network trained on nine faults, and meets eight new ones here
  • Each bar: the share of that fault's samples it flagged as a fault, its recall
  • Fault 18: 92% caught, it looks like the faults it learned
  • Fault 19: 0.1% caught, it looks normal to the network

A classifier only catches new faults that look like the ones it was shown.

Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search

Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, back to Lecture 9

Lecture 9 picked the tree's depth by eye. Its RMSE (root mean squared error) over grouped folds: 9.42 MPa with no depth limit, 9.10 MPa at depth 9.

  • : the hyperparameters, such as the tree depth and the minimum leaf size
  • : the model's parameters, the tree's splits, found by training
  • Every evaluation of the outer problem trains a model: a search needs many, and a record of each
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, Optuna and MLflow

Optuna: an open source Python library that searches hyperparameters for you.

You write an objective function; Optuna proposes the trials and keeps the best. MLflow, from Lectures 1 and 2, records runs. A search uses both:

You write an objective function
→
Optuna proposes a trial, runs it
→
MLflow records it as a run

Every trial stays on record, and the best model is registered.

Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, grid versus random

Two hyperparameters, nine trials for each search. Only the x-axis hyperparameter matters in this example.

0 of 9 trials

The bump (made up for the picture): the validation score against the x-axis hyperparameter. Its peak is the best setting.

Takeaway: same 9 trials, but random search tries 9 values of what matters and grid only 3, so random gets closer to the best.

Bergstra and Bengio (2012), JMLR 13
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, how Optuna picks the next trial

TPE (Tree-structured Parzen Estimator): Optuna's default rule for picking the next trial.

  • The best 10% of the trials so far are good, the rest bad
  • A smoothed histogram of each group: a Parzen estimator
  • Try next where good is common and bad is rare
  • Why TPE: it spends trials where the good results are; random search spends them anywhere

Bergstra et al. (2011), where TPE comes from

Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, Optuna on the concrete strength dataset

Each dot: one trial's validation RMSE. Each line: the best RMSE so far, so it only goes down.

0 of 40 trials

Green dotted lines: Lecture 9's two trees, 9.42 and 9.10 MPa. TPE's best, 8.90 MPa, beats both.

Akiba et al. (2019), the Optuna paper
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, the run hierarchy

A search trains a model once per trial. MLflow keeps every one of them, in one place:

Experimentthis search and all its runs
Parent runthe whole Optuna study
Child run: trial 0its settings, its validation RMSE
Child run: trial 1its settings, its validation RMSE
Child run: trial 2its settings, its validation RMSE
…one child run per trial
Model registrythe best settings, refit and saved by name

Each child run is one trial: a new max_depth and min_samples_leaf, and the validation RMSE they give. The data, the folds and the model stay the same. The registry holds the model to use, with a name and a version anyone can load.

MLflow, hyperparameter tuning with child runs
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, one trial in code

def objective(trial):
    params = dict(
        max_depth=trial.suggest_int('max_depth', 2, 20),
        min_samples_leaf=trial.suggest_int('min_samples_leaf', 1, 50),
    )
    model = DecisionTreeRegressor(
        random_state=SEED,
        **params,
    )
    rmse = -cross_val_score(
        model,
        X_tr,
        y_tr,
        groups=groups_tr,
        cv=GroupKFold(5),
        scoring='neg_root_mean_squared_error',
    ).mean()
    with mlflow.start_run(nested=True):
        mlflow.log_params(params)
        mlflow.log_metric('val_rmse', rmse)
    return rmse
  • def objective(trial): Optuna calls it once per trial. It returns the score to minimize.
  • trial.suggest_int(name, low, high): Optuna picks an integer in that range for this trial.
  • Lecture 9's tree: random_state=SEED fixes its randomness, **params passes this trial's settings.
  • cross_val_score, as in Lecture 9: cv=GroupKFold(5) makes five folds, groups= keeps each mix in one fold, scoring= asks for the RMSE (negative, hence the minus sign).
  • mlflow.start_run(nested=True): a child run inside the parent run, with this trial's settings and score.
  • return rmse: the number Optuna minimizes.
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, the search in code

parent = mlflow.start_run(run_name='concrete-tree-search')
study = optuna.create_study(
    direction='minimize',
    sampler=optuna.samplers.TPESampler(seed=SEED),
)
study.optimize(
    objective,
    n_trials=20,
)

winner = DecisionTreeRegressor(
    random_state=SEED,
    **study.best_params,
).fit(X_tr, y_tr)
mlflow.sklearn.log_model(
    winner,
    name='model',
    registered_model_name='concrete-tree',
    skops_trusted_types=['sklearn.tree._tree.Tree'],
)
mlflow.end_run()
  • mlflow.start_run(run_name=...): opens the parent run. It stays open for the whole search.
  • create_study: direction='minimize', since a lower RMSE is better. sampler= picks TPE, seeded so the search repeats.
  • study.optimize: runs 20 trials, so 20 child runs.
  • The winner: study.best_params, refit on all 835 training rows.
  • log_model: saves and registers the tree as concrete-tree. skops_trusted_types= marks its type as safe to load; MLflow requires it.
  • mlflow.end_run(): closes the parent run.
Systems and Toolchains for AI Engineers
06-763 / L10

Tracking and search, test once

1. Search
on the training mixes only
20 trials, best validation RMSE 8.90 MPa
→
2. Refit and register
the winner, on all 835 training rows
concrete-tree, version 1
→
3. Test once
on the 195 held-out rows
test RMSE 7.41 MPa

Report the test number, 7.41 MPa. The test mixes never took part in the search.

It is lower than 8.90 here: 195 rows from 86 mixes is a small test set, and its score depends on which mixes landed in it.

Systems and Toolchains for AI Engineers
06-763 / L10

Demos

l10-classification.ipynb: the 52-channel classifier, its confusion matrix, three thresholds, and its recall on faults it never saw.

l10-tracking-search.ipynb: a 20-trial Optuna search on the concrete strength dataset, one MLflow child run per trial, the winner registered and tested once.

Classification notebook / Tracking and search notebook

Systems and Toolchains for AI Engineers
06-763 / L10

Recap

Four questions this session added to Lecture 9's four:

What turns a number into a category?

A threshold on a probability.

Fault 4's boundary sits at 43.4% open; fault 14 needs more than one straight cut.

How do you judge a detector?

Its confusion matrix, not accuracy alone.

The baseline scores 85.4% accuracy and 0 recall; precision and recall expose it.

What can a classifier recognize?

New faults only if they look like the ones it was shown.

Two faults it never saw: fault 18 was caught 92% of the time, fault 19 only 0.1%.

How do you trust a search?

Track every trial, test the winner once.

TPE reached 8.90 MPa at trial 11, random search 9.00 at trial 38. The test set gave 7.41.

Report the confusion matrix with its threshold, track every trial, and test once.

Systems and Toolchains for AI Engineers
06-763 / L10

This week

Practice module for this session, for participation credit
Demos l10-classification.ipynb and l10-tracking-search.ipynb, to rerun after class

Systems and Toolchains for AI Engineers

Last time 10, classification 25, measuring a classifier 20, tracking and search 20, demos 10. That is 85 minutes of deck, with 25 left after it.

The plant carries classification and its measures. The concrete strength dataset comes back for the search, with Lecture 9's own decision tree.

Ten minutes before anything new. Same four questions, then capacity again, slower, because I rushed the end of last time.

Same four cards from the end of last time: what a model knows, what training solved, what the score measured, what was holding it back. These apply to the classifier we build today too, so I want them fresh before we start.

Same five knobs as last time, named one at a time: degree, depth, hidden units, alpha, length scale.

Overfit and underfit again, plus the bias-variance line underneath them, since I went past this fast last time. A straight line on the concrete strength dataset is the high-bias end, an unlimited tree the high-variance end.

One figure, one story: the black training curve keeps falling, the red validation curve flattens past depth nine. Looking only at the training score is misleading.

Gap is variance, level is bias. The line's gap closed but its level is still a bit high, so it underfits. The tree's gap never closed. Same four rows as last time, reread slower.

Same fit and predict as Lecture 9. What changes is the output: a probability first, then a threshold turns it into a class.

One channel and one fault first: the valve sits near 41% open in normal runs and near 45% under fault 4. All 52 channels come in later, for the classifier we measure.

Same weighted sum as linear regression. The fraction and the threshold are the new parts. 0.5 is only a starting choice; the threshold slider moves it later.

The boundary is where the weighted sum is zero, so x = -b/w. Normal runs average 41.1% open, fault 4 runs 44.9%, and they barely overlap. The one mistake is a single false alarm.

Stiction: friction holds the valve until the push breaks it free, then it overshoots, and the loop settles into an oscillation. The controller still holds the reactor temperature on average, so the mean stays near 41% open. xmv_10 is the command to the valve, so these swings are the controller fighting the stuck valve.

The normal samples sit in a narrow band, the fault 14 samples spread from about 29% to 54%. A tree of depth 8 catches 99.7%.

Same fit and predict for all four. Each card replays its model learning: logistic regression by gradient descent from zero, the tree one depth at a time, the network iteration by iteration, the Gaussian process as the points arrive. A line, boxes, a smooth curve, smooth probabilities.

A plant runs normally most of the time, so answering "normal" is right most of the time. Accuracy cannot tell this detector from a useful one.

8,300 faults caught, 340 missed, and only 13 false alarms among 50,360 normal samples. Top left and bottom right are right; the other two boxes are the two kinds of mistake.

Precision: when it raises an alarm, can I trust it? Recall: of the real faults, how many did it catch? Both numbers come from the same four boxes.

The baseline is the "always normal" detector from the accuracy slide. Logistic regression misses one fault sample in five; the network misses one in twenty-five.

At 0.01, 8,438 of the 8,640 faults are caught, with 3,664 false alarms. At 0.99 there are no false alarms, and 7,917 are caught. Recall moves little and precision a lot, because this classifier already separates the two classes well.

The blue bar is the nine faults it trained on, 0.961. Nothing in the classifier says in advance which new faults it will catch.

Choosing hyperparameters is an optimization problem with training inside it. The outer problem is what a search solves.

Two jobs: Optuna decides what to try next, MLflow writes down what happened.

Nine trials either way. Grid repeats each value of the important hyperparameter three times, so it tries only 3 values of it. Random almost never repeats a value, so its nine trials cover that axis better.

The first 20 trials of the search on the concrete strength dataset. Good means the best 10%: 2 of 20, both with leaf size 9. The bad ones spread from 2 to 47. Trials 21 to 26 use leaf sizes 8 to 13. Optuna's own curves are smoother than these.

Trials 1 to 10 are the same for both: TPE starts with 10 random trials. From trial 11, TPE tries leaf sizes of 2 to 20, where the good trials were; random keeps drawing 21 to 48. TPE reaches 8.90 at trial 11, random gets to 9.00 only at trial 38.

One run per trial, as in Lectures 1 and 2, now grouped under a parent. Every child run stays on record. The registry is where the chosen model goes.

The same objective as the notebook, one piece at a time. Everything inside it is Lecture 9's grouped cross-validation; the two new parts are suggest_int and the nested run.

The parent run is opened before the study and closed after the registration, so the twenty child runs, the best settings and the registered model all sit under it.

Same rule as Lecture 9, now for the winner of a search: the test set is touched once, after the search is over.

About five minutes each. The first run of the classification notebook downloads 45 MB of plant data.

Each card's picture is the slide where the room saw the answer. The closing line is the three habits under the four cards.