# Assignment 1 — Reproducible Project Scaffold
**Module:** Week 01 · **Released:** Lecture 2 · **Due:** 1 week later

## Goal
Take the broken demo notebook from Lecture 1 and land it in a `uv`-managed project on your own machine that someone else could rebuild and rerun with two commands. You are not writing a new analysis. The analysis already exists, it just cannot be reproduced, and your job is to fix that and give it a home. This scaffold is the template you reuse all semester.

## Learning outcomes assessed
- Create a deterministic environment with `uv`, meaning a pinned interpreter and a committed lockfile.
- Structure a project with a `src/` layout that keeps raw data and the virtual environment out of the code.
- Turn notebook cells into importable functions behind a seeded entry point.
- Log a run to MLflow so a number can be traced back to the code, the data, and the seed that produced it.
- Write a README that tells a stranger how to rebuild and reproduce.

## Starting point
[`l01-reproducibility.ipynb`](../../lectures/l01/l01-reproducibility.ipynb), the notebook from Lecture 1 that does not reproduce. Read it on that page, or get your own copy:

<a href="../../_sources/lectures/l01/l01-reproducibility.ipynb"><strong>Download l01-reproducibility.ipynb</strong></a>

```bash
curl -O https://kitchingroup.cheme.cmu.edu/f26-06763/_sources/lectures/l01/l01-reproducibility.ipynb
```

The same page has a download button and an "Open in Colab" button, if you would rather look at the notebook before committing to a copy. Colab is for reading it: the fixing happens in your own project, where the environment is yours to pin.

It fails in three ways, all of which we diagnosed in class: an absolute path that exists only on the author's machine, an unpinned train/test split that gives a different answer every run, and a NumPy version dependence where exactly one of two integration cells fails depending on which version you installed. Fixing those three is the first half of the work. Putting the fixed analysis somewhere it can survive is the second.

## Dataset
UCI Air Quality Data Set, hourly responses of a metal-oxide multisensor gas device plus reference concentrations, the same data used in both lectures. Your code downloads it or reads it from a documented path.

You are not graded on data wrangling, so the loader is given to you. It handles every quirk the export has, meaning the semicolon separator, the comma as decimal mark, the two empty trailing columns, and the `-200` sentinel that stands in for a missing value:

```python
df = (
    pd.read_csv(path, sep=';', decimal=',')
    .dropna(axis=1, how='all')
    .dropna(how='all')
)
df['ts'] = pd.to_datetime(
    df['Date'] + ' ' + df['Time'].str.replace('.', ':', regex=False),
    format='%d/%m/%Y %H:%M:%S',
)
df = df.replace(-200, np.nan)
```


## Names to use
The script that builds your submission has to find your work, and a grader has to read it. Use these names and both go smoothly:

| Thing | Name |
|---|---|
| Package | `sensorlab`, in `src/sensorlab/` |
| Entry point | `src/sensorlab/train.py`, run as `python -m sensorlab.train --seed 0` |
| Notebook | `notebooks/explore.ipynb` |
| Readme | `README.md` |
| MLflow store | `sqlite:///mlflow.db` |
| Metric | printed to stdout, and logged as `r2` |

If you have already built it under other names, the script goes looking rather than giving up: it finds a package under `src/` or beside `pyproject.toml`, prefers `train.py` and otherwise takes the module that parses arguments and has a `main`, finds any notebook, any `README*`, and whichever MLflow store was written most recently. The report's first page prints what it found. If that is not what you meant, say so directly and run it again:

```bash
python3 a01-evidence.py --andrew-id yourid --module mypkg.fit --notebook analysis.ipynb
```

Two things it cannot work around, so do not get creative with them. Your entry point must accept `--seed`, and it must print the metric where you can see it, because "the same seed prints the same number twice" is checked by running your command and reading the output.

## Tasks
1. **Scaffold the project.** `uv init --package sensorlab`, pin Python 3.11 or newer with `.python-version`, and `uv add pandas scikit-learn matplotlib mlflow` so that `uv.lock` is generated. The `--package` flag is the one that matters: plain `uv init` gives you a loose script rather than an installable package, and `python -m sensorlab.train` then fails with `No module named sensorlab` no matter how correct your `src/` tree looks. Keep the raw data under `data/`, which is where the loader downloads it, and nowhere else in the project. Then delete `.venv`, run `uv sync`, and confirm the project still runs. That deletion is the whole point of the lockfile, so do it once deliberately.
2. **Fix the three defects** in the notebook, keeping it as `notebooks/explore.ipynb`. The path becomes relative, the split gets a fixed seed, and the integration works against whichever NumPy your lockfile pinned. The notebook must run top to bottom after Restart and Run All. Two practical notes: add Jupyter to the project with `uv add --dev jupyter` and launch it with `uv run jupyter lab`, or the kernel will not have your package on its path; and the notebook's working directory is `notebooks/`, so the data path it passes is `../data/AirQualityUCI.csv`.
3. **Refactor into a package.** Move the logic into `src/sensorlab/` as functions, at minimum `load(path)` and `clean(df)`, plus a `train.py` that fits the same baseline model and prints its metric. Give `train.py` an entry point behind `if __name__ == "__main__"` that takes `--seed`, so the analysis runs as `uv run python -m sensorlab.train --seed 0`. The notebook should import these functions rather than keeping a second copy of them.
4. **Track the runs.** Point MLflow at a local SQLite store with `mlflow.set_tracking_uri('sqlite:///mlflow.db')`, log the seed as a parameter and at least one metric, then run the trainer with two different seeds. Open `mlflow ui --backend-store-uri sqlite:///mlflow.db` and confirm both runs are there.
5. **Write the README**, half a page. How to rebuild the environment, how to get the data, and the exact command that reproduces your number. Add one line disclosing any generative-AI use, per the syllabus. It goes into the PDF verbatim, so it is the one part of the submission written in your own voice.

## Snippets you may copy
These exist so your two hours go into the scaffold rather than into syntax you have not been shown yet.

```python
# src/sensorlab/train.py, the entry point
import argparse

def main():
    p = argparse.ArgumentParser()
    p.add_argument('--seed', type=int, default=0)
    p.add_argument('--data-path', default='data/AirQualityUCI.csv')
    args = p.parse_args()
    ...

if __name__ == '__main__':
    main()
```

```python
# logging one run
import mlflow

mlflow.set_tracking_uri('sqlite:///mlflow.db')
mlflow.set_experiment('sensorlab')

with mlflow.start_run(run_name=f'seed-{args.seed}'):
    mlflow.log_params({'seed': args.seed, 'data': Path(args.data_path).name})
    mlflow.log_metric('r2', r2)
```

## What you hand in
**One PDF, uploaded to Canvas.** Everything else stays in your local machine. There is no repository to hand in, nothing to push, and no GitHub account needed. The PDF is generated from your project by a script we provide, which runs the commands below and records what they actually printed, then embeds your `pyproject.toml`, everything in `src/`, your notebook's code cells, and your README, with the code syntax-highlighted so it can be read. Grading reads that one document.

From your project root:

```bash
curl -O https://kitchingroup.cheme.cmu.edu/f26-06763/a01-evidence.py
python3 a01-evidence.py --andrew-id yourid --name "Your Name"
```

It writes `evidence.pdf`. Upload that. There is no print step and nothing to install. Run it with the system `python3` rather than `uv run`, because the first thing it does is delete `.venv` and rebuild it, which is the evidence for correctness.

Read the PDF before you upload it. It prints a PASS or FAIL for each line below, decided from the output rather than asserted, and you get to see exactly what the grader sees. A report with one FAIL and five PASSes is worth more than no report, so a failing check is a reason to fix it and rerun, never a reason not to submit.

The report prints the sha256 of the script that produced it, and the same checksum is published at
<https://kitchingroup.cheme.cmu.edu/f26-06763/a01-evidence.py.sha256>, so run the copy you downloaded rather than a modified one.
It also cross-checks itself: the metric your seed 0 run printed has to turn up again in your MLflow table.

The report's first page prints your score, the fraction of checks that passed, to two decimals. That number is what your TA records, so it is worth seeing it before they do.

## Definition of done
The script decides each of these from your project, and your score is the fraction of them that pass. Run it early, not five minutes before the deadline.

- [ ] The environment rebuilds from the lockfile, after `.venv` is deleted.
- [ ] `python -m sensorlab.train --seed 0`, run twice, prints the same number both times.
- [ ] `--seed 1` prints a different number.
- [ ] The raw data sits under `data/` and nowhere else in the project.
- [ ] The notebook's code cells are numbered 1, 2, 3 and so on, which is what a fresh Restart and Run All leaves behind. Any other numbering is the hidden-state problem from Lecture 1, showing in your own submission.
- [ ] The metric your seed 0 run printed also appears in the MLflow table, so the number on the page and the number in the store are the same number.

## How this is graded
Your score is the fraction of the Definition-of-done checks that pass, printed on the report's first page to two decimals, decided by the script from your project rather than from your description of it. Your TA also reads what the report embeds, your package, your notebook's code cells and your README, and can adjust from there when something is clearly better or worse than the checks can see. Nothing outside the PDF is graded, so anything you want considered has to be in it.

A check that fails costs you that check and nothing else. The report is built so that one mistake stays one mistake: a rejected lockfile, for instance, does not also take your runs down with it.

## Allowed tools & AI-use note
Generative-AI assistance is allowed with disclosure per the syllabus. Note in your README what you used it for, and remember that the README is embedded in the PDF you submit. Editing the generated report by hand is falsifying a submission, and it is also more work than rerunning the script. You must be able to explain every part of your submission on request, including why the lockfile matters and what each function does. Standard library plus the packages listed above unless you document additions.

## Stretch (optional, not graded)
None of these are required, and none of them are covered in Week 1. They are here for anyone who wants to keep going.

- Add a `pytest` test, for example that `clean` leaves no `-200` values behind.
- Log the figure as an MLflow artifact alongside the metric.
- Add a `Makefile` or `justfile` wrapping `setup`, `run`, and `test`.
- Put the project under version control: `git init`, a `.gitignore` covering `data/`, `.venv/`, `mlruns/` and `mlflow.db`, and a first commit. Assignment 2 assumes you have done this at least once.
