Provenance: the record of which data, which code, and which parameters produced a given number.
If the answer is "a notebook run out of order
on a laptop that has been reimaged,"
there is no answer.
Control function commanding nose-down stabilizer trim.
One bad sensor was sufficient.
Left AoA sensor biased by ~21°
Traced to a replacement sensor
mis-calibrated during an earlier repair
A measurement channel nobody had validated,
driving a flight control surface.
Boeing's hazard assessment evaluated
erroneous data from both air data channels
→ "beyond extremely improbable"
But the MCAS path was exposed to a single failure.
House T&I Committee report, p. 106
Incorrect assumptions about flight crew response,
plus incomplete review of flight deck effects,
are what made single-sensor reliance appear acceptable.
JT610, Oct 2018: 189 killed.
ET302, Mar 2019: 157 killed.
If (3) is "an operator will compensate,"
that needs evidence, not assumption.
A CSV has no schema, no types, no constraints,
and no way to read one column without reading all of them.
A spreadsheet datastore is the far end of that road.
| Model | Example | Good at |
|---|---|---|
| Relational | PostgreSQL | continuous writes, joins, correctness |
| Columnar file | Parquet | scanning few columns of many |
| Embedded analytical | DuckDB | SQL over Parquet, no server |
A fourth, the vector store, indexes by similarity.
Not "which database is best" but
"what is my access pattern?"
Batch runs on a schedule over bounded data. Most engineering work needs only this.
Streaming processes records as they arrive: buys latency, costs correctness.
A pipeline that silently passes bad data
is worse than one that crashes.
Declare what you expect:
column types, physical ranges, null rates, row counts
Fail loudly when reality disagrees.
It is an optimization loop.
Automatic differentiation: frameworks record the operations you perform, then replay them backwards for exact gradients.
You never hand-derive anything. That is the whole trick behind PyTorch and JAX.
Where there is structure to exploit:
But: a boosted tree on good features beats
a neural net on bad ones, most of the time.
Beat the baseline first.
Both failures today were evaluation failures.
After a hundred runs, "which config produced this?"
is unanswerable from memory. One run = one reproducible fact.
Making it run somewhere else, repeatedly,
for someone who is not you.
Container: an image bundling code, dependencies and system libraries into one artifact that runs anywhere the runtime exists.
A VM virtualizes hardware and boots an OS; a container shares the host kernel
and isolates the process. Milliseconds, not tens of seconds.
FastAPI: your model as an HTTP endpoint,
with typed request and response schemas.
That schema is also a data contract with your callers.
Then the questions turn operational:
latency budget, throughput, cost per prediction, behavior under load
Most courses stop at deployment.
Models degrade because the world moves,
exactly as the calibration figure showed.
Watch: input distributions, prediction distributions,
the gap against ground truth once labels arrive
MLOps automates that loop: CI that tests data and models and not only
code, reproducible retraining, staged rollout so a bad model does not reach
everyone, and the ability to roll back.
Flu Trends is what absence looks like.
A next-token predictor trained on a very large corpus.
What matters operationally:
Two integration patterns:
For engineering: agents over your database,
your simulation, your instrument.
| Layer | Tool | Prevents |
|---|---|---|
| Environments | uv |
merely-probable rebuilds |
| Storage | Postgres/DuckDB/Parquet | CSV "mess", truncation |
| Dataframes | pandas/Polars | OOM (out-of-memory), unreadable transforms |
| Validation | pandera | "bad data" passing "quietly" |
| Tracking | MLflow | unattributable results |
| ML | PyTorch/JAX | hand-derived gradients |
| Serving | FastAPI/Docker | "works on my machine" |
LLM and agent frameworks are deliberately not standardized:
that ecosystem changes much faster than a semester, so we teach interfaces
and evaluation techniques rather than vendor's (OpenAI, Anthropic) specifics.
l01-reproducibility.ipynbUCI Air Quality → plot → naive split → fresh checkout
Diagnose each break before I do.
Compare failures with your neighbor on 3.
None of the three announces itself while you are making it.
Install before next class git, Python 3.11+, uv
Assignment 1 released next session
Practice module ~10 min, ends in a PDF you upload: https://kitchingroup.cheme.cmu.edu/f26-06763/game/#/l01
Full notes, with all sources on the course website.
110 minutes. Budget: 10 / 30 / 20 / 35 / 10, leaving ~5 slack. The tour is the part students actually need today. If you are running long, cut MCAS detail, not the tour.
Worth 60 seconds. Sets the tone for citation discipline all semester.