Assignment 1 — Reproducible Project Scaffold#
Module: Week 01 · Released: Lecture 2 · Due: 1 week later
Goal#
Take the broken demo notebook from Lecture 1 and land it in a uv-managed project on your own machine that someone else could rebuild and rerun with two commands. You are not writing a new analysis. The analysis already exists, it just cannot be reproduced, and your job is to fix that and give it a home. This scaffold is the template you reuse all semester.
Learning outcomes assessed#
Create a deterministic environment with
uv, meaning a pinned interpreter and a committed lockfile.Structure a project with a
src/layout that keeps raw data and the virtual environment out of the code.Turn notebook cells into importable functions behind a seeded entry point.
Log a run to MLflow so a number can be traced back to the code, the data, and the seed that produced it.
Write a README that tells a stranger how to rebuild and reproduce.
Starting point#
l01-reproducibility.ipynb, the notebook from Lecture 1 that does not reproduce. Read it on that page, or get your own copy:
Download l01-reproducibility.ipynb
curl -O https://kitchingroup.cheme.cmu.edu/f26-06763/_sources/lectures/l01/l01-reproducibility.ipynb
The same page has a download button and an “Open in Colab” button, if you would rather look at the notebook before committing to a copy. Colab is for reading it: the fixing happens in your own project, where the environment is yours to pin.
It fails in three ways, all of which we diagnosed in class: an absolute path that exists only on the author’s machine, an unpinned train/test split that gives a different answer every run, and a NumPy version dependence where exactly one of two integration cells fails depending on which version you installed. Fixing those three is the first half of the work. Putting the fixed analysis somewhere it can survive is the second.
Dataset#
UCI Air Quality Data Set, hourly responses of a metal-oxide multisensor gas device plus reference concentrations, the same data used in both lectures. Your code downloads it or reads it from a documented path.
You are not graded on data wrangling, so the loader is given to you. It handles every quirk the export has, meaning the semicolon separator, the comma as decimal mark, the two empty trailing columns, and the -200 sentinel that stands in for a missing value:
df = (
pd.read_csv(path, sep=';', decimal=',')
.dropna(axis=1, how='all')
.dropna(how='all')
)
df['ts'] = pd.to_datetime(
df['Date'] + ' ' + df['Time'].str.replace('.', ':', regex=False),
format='%d/%m/%Y %H:%M:%S',
)
df = df.replace(-200, np.nan)
Names to use#
The script that builds your submission has to find your work, and a grader has to read it. Use these names and both go smoothly:
Thing |
Name |
|---|---|
Package |
|
Entry point |
|
Notebook |
|
Readme |
|
MLflow store |
|
Metric |
printed to stdout, and logged as |
If you have already built it under other names, the script goes looking rather than giving up: it finds a package under src/ or beside pyproject.toml, prefers train.py and otherwise takes the module that parses arguments and has a main, finds any notebook, any README*, and whichever MLflow store was written most recently. The report’s first page prints what it found. If that is not what you meant, say so directly and run it again:
python3 a01-evidence.py --andrew-id yourid --module mypkg.fit --notebook analysis.ipynb
Two things it cannot work around, so do not get creative with them. Your entry point must accept --seed, and it must print the metric where you can see it, because “the same seed prints the same number twice” is checked by running your command and reading the output.
Tasks#
Scaffold the project.
uv init --package sensorlab, pin Python 3.11 or newer with.python-version, anduv add pandas scikit-learn matplotlib mlflowso thatuv.lockis generated. The--packageflag is the one that matters: plainuv initgives you a loose script rather than an installable package, andpython -m sensorlab.trainthen fails withNo module named sensorlabno matter how correct yoursrc/tree looks. Keep the raw data underdata/, which is where the loader downloads it, and nowhere else in the project. Then delete.venv, runuv sync, and confirm the project still runs. That deletion is the whole point of the lockfile, so do it once deliberately.Fix the three defects in the notebook, keeping it as
notebooks/explore.ipynb. The path becomes relative, the split gets a fixed seed, and the integration works against whichever NumPy your lockfile pinned. The notebook must run top to bottom after Restart and Run All. Two practical notes: add Jupyter to the project withuv add --dev jupyterand launch it withuv run jupyter lab, or the kernel will not have your package on its path; and the notebook’s working directory isnotebooks/, so the data path it passes is../data/AirQualityUCI.csv.Refactor into a package. Move the logic into
src/sensorlab/as functions, at minimumload(path)andclean(df), plus atrain.pythat fits the same baseline model and prints its metric. Givetrain.pyan entry point behindif __name__ == "__main__"that takes--seed, so the analysis runs asuv run python -m sensorlab.train --seed 0. The notebook should import these functions rather than keeping a second copy of them.Track the runs. Point MLflow at a local SQLite store with
mlflow.set_tracking_uri('sqlite:///mlflow.db'), log the seed as a parameter and at least one metric, then run the trainer with two different seeds. Openmlflow ui --backend-store-uri sqlite:///mlflow.dband confirm both runs are there.Write the README, half a page. How to rebuild the environment, how to get the data, and the exact command that reproduces your number. Add one line disclosing any generative-AI use, per the syllabus. It goes into the PDF verbatim, so it is the one part of the submission written in your own voice.
Snippets you may copy#
These exist so your two hours go into the scaffold rather than into syntax you have not been shown yet.
# src/sensorlab/train.py, the entry point
import argparse
def main():
p = argparse.ArgumentParser()
p.add_argument('--seed', type=int, default=0)
p.add_argument('--data-path', default='data/AirQualityUCI.csv')
args = p.parse_args()
...
if __name__ == '__main__':
main()
# logging one run
import mlflow
mlflow.set_tracking_uri('sqlite:///mlflow.db')
mlflow.set_experiment('sensorlab')
with mlflow.start_run(run_name=f'seed-{args.seed}'):
mlflow.log_params({'seed': args.seed, 'data': Path(args.data_path).name})
mlflow.log_metric('r2', r2)
What you hand in#
One PDF, uploaded to Canvas. Everything else stays in your local machine. There is no repository to hand in, nothing to push, and no GitHub account needed. The PDF is generated from your project by a script we provide, which runs the commands below and records what they actually printed, then embeds your pyproject.toml, everything in src/, your notebook’s code cells, and your README, with the code syntax-highlighted so it can be read. Grading reads that one document.
From your project root:
curl -O https://kitchingroup.cheme.cmu.edu/f26-06763/a01-evidence.py
python3 a01-evidence.py --andrew-id yourid --name "Your Name"
It writes evidence.pdf. Upload that. There is no print step and nothing to install. Run it with the system python3 rather than uv run, because the first thing it does is delete .venv and rebuild it, which is the evidence for correctness.
Read the PDF before you upload it. It prints a PASS or FAIL for each line below, decided from the output rather than asserted, and you get to see exactly what the grader sees. A report with one FAIL and five PASSes is worth more than no report, so a failing check is a reason to fix it and rerun, never a reason not to submit.
The report prints the sha256 of the script that produced it, and the same checksum is published at https://kitchingroup.cheme.cmu.edu/f26-06763/a01-evidence.py.sha256, so run the copy you downloaded rather than a modified one. It also cross-checks itself: the metric your seed 0 run printed has to turn up again in your MLflow table.
The report’s first page prints your score, the fraction of checks that passed, to two decimals. That number is what your TA records, so it is worth seeing it before they do.
Definition of done#
The script decides each of these from your project, and your score is the fraction of them that pass. Run it early, not five minutes before the deadline.
[ ] The environment rebuilds from the lockfile, after
.venvis deleted.[ ]
python -m sensorlab.train --seed 0, run twice, prints the same number both times.[ ]
--seed 1prints a different number.[ ] The raw data sits under
data/and nowhere else in the project.[ ] The notebook’s code cells are numbered 1, 2, 3 and so on, which is what a fresh Restart and Run All leaves behind. Any other numbering is the hidden-state problem from Lecture 1, showing in your own submission.
[ ] The metric your seed 0 run printed also appears in the MLflow table, so the number on the page and the number in the store are the same number.
How this is graded#
Your score is the fraction of the Definition-of-done checks that pass, printed on the report’s first page to two decimals, decided by the script from your project rather than from your description of it. Your TA also reads what the report embeds, your package, your notebook’s code cells and your README, and can adjust from there when something is clearly better or worse than the checks can see. Nothing outside the PDF is graded, so anything you want considered has to be in it.
A check that fails costs you that check and nothing else. The report is built so that one mistake stays one mistake: a rejected lockfile, for instance, does not also take your runs down with it.
Allowed tools & AI-use note#
Generative-AI assistance is allowed with disclosure per the syllabus. Note in your README what you used it for, and remember that the README is embedded in the PDF you submit. Editing the generated report by hand is falsifying a submission, and it is also more work than rerunning the script. You must be able to explain every part of your submission on request, including why the lockfile matters and what each function does. Standard library plus the packages listed above unless you document additions.
Stretch (optional, not graded)#
None of these are required, and none of them are covered in Week 1. They are here for anyone who wants to keep going.
Add a
pytesttest, for example thatcleanleaves no-200values behind.Log the figure as an MLflow artifact alongside the metric.
Add a
Makefileorjustfilewrappingsetup,run, andtest.Put the project under version control:
git init, a.gitignorecoveringdata/,.venv/,mlruns/andmlflow.db, and a first commit. Assignment 2 assumes you have done this at least once.