L2 demo: from an empty directory to a tracked run#

In L1 an analysis failed to reproduce three ways: an absolute path, an unpinned split, and an unrecorded library version. Here we rebuild the same analysis the way it should have been built the first time: a uv-managed project with a locked environment, a pinned interpreter, relative paths, a seeded model, and each run logged to MLflow.

There are two kinds of cell below. The scaffold steps are terminal commands, shown in fenced blocks; run those in a terminal, not in the notebook. The Python cells define the functions that, in the real project, live in src/sensorlab/, and then run them, so this notebook still executes top to bottom on its own.

Data: UCI Air Quality Data Set, the same roadside sensor series from L1, fetched from UCI on first run.

Run this first on Colab#

Colab starts from its own preinstalled environment rather than this course’s uv environment, so run the cell below before anything else. It installs what this notebook needs and Colab does not already have. Outside Colab it does nothing, so you can run it or skip it.

This project is built locally with uv, which is the whole point of the lecture. If you are working locally, skip this cell: it exists only for reading the notebook through the Open-in-Colab button, and does nothing otherwise.

# Run this first on Colab. Anywhere else this cell does nothing.
#
# Only genuinely missing packages are installed, so Colab's own versions of
# everything it already ships are left alone.
import importlib.util
import subprocess
import sys

REQUIREMENTS = {
    "mlflow": "mlflow",
    "numpy": "numpy",
    "pandas": "pandas",
    "sklearn": "scikit-learn",
}


def _missing(module):
    try:
        return importlib.util.find_spec(module) is None
    except ModuleNotFoundError:  # the parent package is absent
        return True


if "google.colab" in sys.modules:
    need = sorted({pip for mod, pip in REQUIREMENTS.items() if _missing(mod)})
    if need:
        print("installing:", " ".join(need))
        subprocess.run([sys.executable, "-m", "pip", "install", "-q", *need], check=True)
    print("Colab setup done." if need else "Colab: nothing to install.")

1. Scaffold the project (in a terminal)#

uv init --package sensorlab && cd sensorlab
uv add pandas scikit-learn mlflow

uv init --package writes pyproject.toml, .python-version, and an importable src/sensorlab/ (the --package flag is what makes python -m sensorlab.train work); uv add resolves the whole dependency graph into uv.lock. On any other machine, uv sync rebuilds this exact environment. That lockfile, not a requirements.txt, is what makes the rebuild deterministic rather than merely probable.

The layout the project grows into:

sensorlab/
├── pyproject.toml
├── uv.lock
├── .python-version
├── .gitignore          # data/, .venv/, mlflow.db
├── src/sensorlab/      # load, clean, featurize, train
├── data/               # git-ignored: raw data does not go in git
└── tests/

2. The functions (src/sensorlab/)#

In the project these live in importable modules; here we define them in the notebook so the whole thing runs top to bottom. Two of the L1 defects are fixed in the very first function: the data path is relative and its parent is created if missing, so the fetch works on any machine rather than only the author’s.

import io
import hashlib
import urllib.request
import zipfile
from pathlib import Path

import numpy as np
import pandas as pd

DATA = Path('data/AirQualityUCI.csv')
URL = 'https://archive.ics.uci.edu/static/public/360/air+quality.zip'


def load(path: Path = DATA) -> pd.DataFrame:
    """Fetch (once) and parse the UCI Air Quality CSV.

    Relative path, parent created if missing: the L1 bug, fixed. The export is
    semicolon separated with comma decimals, has two trailing empty columns, and
    codes missing values as -200.
    """
    path.parent.mkdir(parents=True, exist_ok=True)
    if not path.exists():
        print(f'fetching {URL}')
        with urllib.request.urlopen(URL) as response:
            payload = response.read()
        with zipfile.ZipFile(io.BytesIO(payload)) as archive:
            path.write_bytes(archive.read('AirQualityUCI.csv'))
    df = (
        pd.read_csv(path, sep=';', decimal=',')
        .dropna(axis=1, how='all')
        .dropna(how='all')
    )
    df['ts'] = pd.to_datetime(
        df['Date'] + ' ' + df['Time'].str.replace('.', ':', regex=False),
        format='%d/%m/%Y %H:%M:%S',
    )
    return df.replace(-200, np.nan)


raw = load()
print(f'{len(raw)} rows, {raw.ts.min().date()} to {raw.ts.max().date()}')

clean and featurize take a frame and return values, with no hidden global state, so they can be imported and tested in isolation. We predict the reference CO measurement from the cheap sensor channels plus temperature and humidity.

FEATURES = [
    'PT08.S1(CO)', 'PT08.S2(NMHC)', 'PT08.S3(NOx)',
    'PT08.S4(NO2)', 'PT08.S5(O3)', 'T', 'RH', 'AH',
]
TARGET = 'CO(GT)'


def clean(df: pd.DataFrame) -> pd.DataFrame:
    """Keep rows with all features and the reference present, in time order."""
    return (
        df.dropna(subset=FEATURES + [TARGET])
        .sort_values('ts')
        .reset_index(drop=True)
    )


def featurize(df: pd.DataFrame):
    """Return the feature matrix and target as arrays."""
    return df[FEATURES].to_numpy(), df[TARGET].to_numpy()


clean_df = clean(raw)
X, y = featurize(clean_df)
print(f'{len(clean_df)} usable rows, {X.shape[1]} features')

train takes a seed. The split is temporal (fit on the earlier 75%, test on the later 25%), which is the honest protocol for a sensor series, and the model is seeded so the run is reproducible.

from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import r2_score


def train(X, y, seed: int) -> float:
    """Temporal split, seeded model, return R2 on the held-out later period."""
    cut = int(len(X) * 0.75)
    model = RandomForestRegressor(n_estimators=200, random_state=seed)
    model.fit(X[:cut], y[:cut])
    return r2_score(y[cut:], model.predict(X[cut:]))


print(f'R2 (seed 0): {train(X, y, seed=0):.4f}')

3. Track each run (MLflow)#

Now make each run a fact you can point to. We log the seed, a hash of the data, and the git commit, alongside the metric, then run the trainer twice with two seeds. MLflow stores the runs locally in a small SQLite file, with no server to run.

import subprocess

import mlflow

# MLflow 3 recommends a local database backend over the bare file store.
mlflow.set_tracking_uri('sqlite:///mlflow.db')
mlflow.set_experiment('l02-scaffold')


def data_hash(path: Path) -> str:
    return hashlib.sha256(path.read_bytes()).hexdigest()[:12]


def git_sha() -> str:
    try:
        out = subprocess.check_output(
            ['git', 'rev-parse', '--short', 'HEAD'], text=True, stderr=subprocess.DEVNULL)
        return out.strip()
    except Exception:
        return 'unknown'


for seed in (0, 1):
    with mlflow.start_run(run_name=f'seed-{seed}'):
        r2 = train(X, y, seed=seed)
        mlflow.log_params({
            'seed': seed,
            'model': 'rf200',
            'data_sha': data_hash(DATA),
            'git_sha': git_sha(),
        })
        mlflow.log_metric('r2', r2)
        print(f'seed {seed}: R2 = {r2:.4f}')

4. Move it into the project#

You do not keep this in the notebook. You move the functions above, unchanged, into src/sensorlab/train.py. One thing has to be added that the notebook never needed: a command-line entry point, so the analysis runs as a command you type rather than cells you click in the right order.

Two things get added at the bottom of train.py. First, wrap the run-and-log logic (the for-loop from the MLflow cell above) into a main(seed) that does one seed. Second, add the command-line entry point that reads --seed and calls it. Both use the functions and DATA you already moved in (load, clean, featurize, train, data_hash, git_sha):

def main(seed: int) -> None:
    mlflow.set_tracking_uri("sqlite:///mlflow.db")
    mlflow.set_experiment("l02-scaffold")
    X, y = featurize(clean(load()))
    with mlflow.start_run(run_name=f"seed-{seed}"):
        r2 = train(X, y, seed=seed)
        mlflow.log_params({"seed": seed, "model": "rf200",
                           "data_sha": data_hash(DATA), "git_sha": git_sha()})
        mlflow.log_metric("r2", r2)
        print(f"seed {seed}: R2 = {r2:.4f}")


if __name__ == "__main__":
    import argparse

    parser = argparse.ArgumentParser()
    parser.add_argument("--seed", type=int, default=0)
    args = parser.parse_args()
    main(args.seed)

Why it is needed, and what each part does:

  • if __name__ == "__main__": runs the block only when the file is executed directly (python -m sensorlab.train), and not when it is imported (from sensorlab.train import train). That is what lets one file be both an importable library and a runnable program. Without the guard, this code would fire every time anything imported the module, including your notebook and your tests.

  • argparse is Python’s standard command-line reader. add_argument("--seed", ...) declares a --seed option; parse_args() reads what you actually typed.

  • main(args.seed) runs one seeded, logged run. The seed becomes a knob you set on the command line instead of a number buried in a cell, which is the whole point: the run is explicit and repeatable.

It is shown here, not run: parse_args() inside a notebook would try to parse Jupyter’s own launch arguments and fail. In the project you run it from a terminal, once per seed:

uv run python -m sensorlab.train --seed 0
uv run python -m sensorlab.train --seed 1

5. Compare the two runs#

The two runs are now in mlflow.db, whether you logged them here (the MLflow cell above) or from the terminal (the two commands above). Open the UI to compare them.

This is a local command: it serves to http://127.0.0.1:5000 on your own machine, so run it from the project directory. On Colab there is nothing for it to open.

uv run mlflow ui --backend-store-uri sqlite:///mlflow.db   # then open http://127.0.0.1:5000

You will see two runs that differ only by their seed, with slightly different R2. That small difference is the whole reason to log the seed: without it, neither number is reconstructible. The data hash and git SHA make the rest of the run reconstructible too, which is the provenance the L2 notes argue for.