L10 demo: a hyperparameter search, tracked and registered#
Optuna is a Python library that runs a hyperparameter search for you: https://optuna.org. You give it a range for each hyperparameter and a way to score one choice, and it tries a sequence of configurations and keeps the best one.
MLflow records every trial Optuna tries as its own run, so the settings and their scores survive after the search ends.
Lecture 9 fit its decision tree to the concrete strength dataset by trying a few depths by eye. This notebook lets Optuna search over the tree’s depth and leaf size instead. We then register the best tree, load it back by name, and touch the held-out test mixes exactly once.
Companion notes:
notes.md.
Run this first on Colab#
Colab starts from its own preinstalled environment rather than this course’s uv
environment, so run the cell below before anything else. It installs what this
notebook needs and Colab does not already have. Outside Colab it does nothing, so
you can run it or skip it.
# Run this first on Colab. Anywhere else this cell does nothing.
#
# Only genuinely missing packages are installed, so Colab's own versions of
# everything it already ships are left alone.
import importlib.util
import subprocess
import sys
REQUIREMENTS = {
"mlflow": "mlflow",
"optuna": "optuna",
"pandas": "pandas",
"sklearn": "scikit-learn",
"xlrd": "xlrd",
}
def _missing(module):
try:
return importlib.util.find_spec(module) is None
except ModuleNotFoundError: # the parent package is absent
return True
if "google.colab" in sys.modules:
need = sorted({pip for mod, pip in REQUIREMENTS.items() if _missing(mod)})
if need:
print("installing:", " ".join(need))
subprocess.run([sys.executable, "-m", "pip", "install", "-q", *need], check=True)
print("Colab setup done." if need else "Colab: nothing to install.")
1. Load the concrete strength dataset, and lock the split#
One row per specimen, 20% of the mixes held out exactly as in Lecture 9:
GroupShuffleSplit keeps every row of a mix on the same side, so the model is never
tested on a recipe it trained on.
import io
import urllib.request
import zipfile
from pathlib import Path
import mlflow
import optuna
import pandas as pd
from sklearn.tree import DecisionTreeRegressor
from sklearn.metrics import root_mean_squared_error
from sklearn.model_selection import GroupKFold, GroupShuffleSplit, cross_val_score
DATA = Path('data')
DATA.mkdir(exist_ok=True)
XLS = DATA / 'Concrete_Data.xls'
if not XLS.exists():
url = 'https://archive.ics.uci.edu/static/public/165/concrete+compressive+strength.zip'
XLS.write_bytes(
zipfile.ZipFile(io.BytesIO(urllib.request.urlopen(url).read())).read('Concrete_Data.xls')
)
COLUMNS = ['cement', 'slag', 'fly_ash', 'water', 'superplasticizer',
'coarse_agg', 'fine_agg', 'age_days', 'strength_mpa']
FEATURES, MIX = COLUMNS[:8], COLUMNS[:7]
concrete = pd.read_excel(XLS)
concrete.columns = COLUMNS
X = concrete[FEATURES].to_numpy()
y = concrete['strength_mpa'].to_numpy()
groups = concrete.groupby(MIX).ngroup().to_numpy()
splitter = GroupShuffleSplit(
n_splits=1,
test_size=0.2,
random_state=42,
)
train, test = next(splitter.split(X, y, groups))
X_tr, y_tr, groups_tr = X[train], y[train], groups[train]
X_te, y_te = X[test], y[test]
print(f'{len(X_tr)} training rows ({len(set(groups_tr))} mixes), {len(X_te)} held-out test rows')
835 training rows (342 mixes), 195 held-out test rows
2. Point MLflow at a local store#
The same pattern as Lecture 1 and Lecture 2: a local SQLite file, no server to run. Every trial the search tries becomes one MLflow run, and Optuna’s study becomes the parent run those trials nest under.
SEED = 0
mlflow.set_tracking_uri('sqlite:///mlflow.db')
# An experiment deleted in the MLflow UI keeps its name reserved, so bring it back.
client = mlflow.MlflowClient()
old = client.get_experiment_by_name('concrete-search')
if old is not None and old.lifecycle_stage == 'deleted':
client.restore_experiment(old.experiment_id)
mlflow.set_experiment('concrete-search')
optuna.logging.set_verbosity(optuna.logging.WARNING)
2026/09/28 11:03:37 INFO mlflow.store.db.utils: Creating initial MLflow database tables...
2026/09/28 11:03:37 INFO mlflow.store.db.utils: Updating database tables
2026/09/28 11:03:38 INFO mlflow.tracking.fluent: Experiment with name 'concrete-search' does not exist. Creating a new experiment.
Open the MLflow UI now, so you can watch the search fill it in. In a terminal, from this notebook’s folder:
mlflow ui --backend-store-uri sqlite:///mlflow.db
Start it in this folder: a UI started anywhere else opens an empty mlflow.db there and
shows nothing. Then open http://127.0.0.1:5000 and click the concrete-search experiment.
It is empty
for now. Each trial in section 4 appears as a child run under the parent run as soon as it
finishes (refresh the page), and section 5 adds the registered model under Models.
If the UI crashes with cannot import name 'Traversable', your MLflow is too old for
Python 3.14: run pip install -U mlflow (3.16 works). On Colab there is no terminal for
the UI; mlflow.search_runs() lists the same runs as a table.
3. One trial, one nested run#
objective is what Optuna calls once per trial. It draws a tree depth and a leaf size
from a range, scores that one tree by five-fold GroupKFold on the training mixes, and
logs the trial as a run nested under whichever run is currently open.
max_depth caps how many splits deep the tree can go. min_samples_leaf sets the
fewest training rows a leaf is allowed to end with, so the tree cannot keep splitting
down to one row.
def objective(trial):
params = dict(
max_depth=trial.suggest_int('max_depth', 2, 20),
min_samples_leaf=trial.suggest_int('min_samples_leaf', 1, 50),
)
model = DecisionTreeRegressor(
random_state=SEED,
**params,
)
rmse = -cross_val_score(
model,
X_tr,
y_tr,
groups=groups_tr,
cv=GroupKFold(5),
scoring='neg_root_mean_squared_error',
).mean()
with mlflow.start_run(nested=True):
mlflow.log_params(params)
mlflow.log_metric('val_rmse', rmse)
return rmse
4. Run the study#
mlflow.start_run opens the parent and leaves it open, so every nested run that
objective starts over the next twenty trials lands underneath it.
The sampler is TPE, the Tree-structured Parzen Estimator: Optuna’s default way to choose
the next trial. It looks at the trials so far and tries next where the good ones cluster.
seed=SEED fixes its random draws, so the search gives the same trials every run.
Once the study is done, its best parameters and best score are logged straight to that same parent run.
parent = mlflow.start_run(run_name='concrete-tree-search')
study = optuna.create_study(
direction='minimize',
sampler=optuna.samplers.TPESampler(seed=SEED),
)
study.optimize(
objective,
n_trials=20,
)
mlflow.log_params(study.best_params)
mlflow.log_metric('best_val_rmse', study.best_value)
print(f'best validation RMSE {study.best_value:.2f} MPa after {len(study.trials)} trials')
print('best params:', study.best_params)
best validation RMSE 8.90 MPa after 20 trials
best params: {'max_depth': 9, 'min_samples_leaf': 9}
5. Register the winner#
Refit the best settings on all 835 training rows and register the tree under a name,
concrete-tree. MLflow needs one extra argument to save a decision tree, explained in the
code comment. The parent run then closes, holding twenty child runs, the best settings
and the registered model.
winner = DecisionTreeRegressor(
random_state=SEED,
**study.best_params,
).fit(X_tr, y_tr)
mlflow.sklearn.log_model(
winner,
name='model',
registered_model_name='concrete-tree',
# A DecisionTreeRegressor stores its splits in a type MLflow's default
# serializer does not trust automatically. We just trained this model
# ourselves, so trusting it back is safe.
skops_trusted_types=['sklearn.tree._tree.Tree'],
)
mlflow.end_run()
print('registered concrete-tree; parent run closed')
registered concrete-tree; parent run closed
Successfully registered model 'concrete-tree'.
Created version '1' of model 'concrete-tree'.
6. Load it back, and test once#
A registered model is loaded fresh by its models:/ URI, the way a teammate on a
different machine would load it. This cell is the only one that touches the held-out
test mixes, and it touches them once.
from mlflow import MlflowClient
client = MlflowClient()
latest = max(int(v.version) for v in client.search_model_versions("name='concrete-tree'"))
loaded = mlflow.sklearn.load_model(f'models:/concrete-tree/{latest}')
test_rmse = root_mean_squared_error(y_te, loaded.predict(X_te))
print(f'loaded models:/concrete-tree/{latest}')
print(f'best VALIDATION RMSE (search optimized) : {study.best_value:.2f} MPa')
print(f'held-out TEST RMSE (reported once) : {test_rmse:.2f} MPa')
loaded models:/concrete-tree/1
best VALIDATION RMSE (search optimized) : 8.90 MPa
held-out TEST RMSE (reported once) : 7.41 MPa
Try it#
In the MLflow UI you opened in section 2, sort the child runs of concrete-search by
val_rmse and see which of the two hyperparameters moves the score. Then raise n_trials
in section 4 from 20 to 50, Restart and Run All, and watch the new parent run gain more
child runs.