Systems and Toolchains for AI Engineers
Lecture 3 gave you SQL: describe the result, the planner works out how. Lecture 4 gave you a faster layout for that same idea, still inside a database.
So why leave the database? Because this lecture's work is awkward to write as a query:
That is ordinary programming. Python is better at it.
The Tennessee Eastman process: a simulated chemical plant.
What we want out of it: one row per fault, showing what each sensor read on average while that fault was running. Put two faults side by side and you can see which sensors moved.
Tennessee Eastman data, CC0 · Downs and Vogel, 1993
Reactor, condenser, compressor, separator, stripper: the plant behind this session's 52 columns. Lyu, Botcha, Kulkarni, Pagaria, Alves, Sunshine, and Kitchin (2026). Run it yourself: TEP Studio
Leaving the database costs you something. Write this as one long Python script and two things it was handling for you become your problem:
Polars gets the speed back. Small, restartable stages get the safety back.
A dataframe is a table in memory: named columns, each with a fixed type. You have used one since Lecture 3.
You already wrote grouping and joining in SQL. Here they are methods, not clauses.
readings.group_by("faultNumber").agg(pl.col("^xmeas_.*$").mean()) # every measured column
group_by
GROUP BY
faultNumber = 4
pandas, group by
This one does not carry over. SQL never let you handle rows one at a time: you asked for a result, and the database walked the rows however it wanted.
Vectorization: compute on a whole column at once instead of one value at a time.
On the reactor-pressure column, 480,000 readings:
readings["xmeas_7"] > threshold
for
iterrows()
Hundreds of times slower, and that is before anyone reaches for a bigger machine.
pandas runs each line the moment you write it, on one core. That is eager execution.
Polars: a table library like pandas, but built to use every core at once. If you ask it to, it will also plan the whole computation before running any of it.
Lecture 3: you describe the result in SQL, the query planner decides the mechanism. Same idea, now in Python.
Lazy evaluation: writing an instruction does not run it. It adds a step to a plan. Nothing touches the data until .collect(), which runs the whole plan in one pass.
.collect()
pipeline = (pl.scan_parquet("data/tep.parquet") # reads nothing yet .group_by("faultNumber") .agg(pl.col("^xmeas_.*$").mean())) result = pipeline.collect() # optimize, then run
You met these in Lecture 4, when DuckDB applied them to Parquet:
Because pandas and Polars both lay columns out the same way in memory (Arrow), converting between them is cheap too: prototype in whichever you know, convert only if it turns out to matter.
Polars, Lazy API · Apache Arrow
A pipeline written as one long function has no seam: no smaller piece you can test on its own, and none you can re-run on its own. Break it into stages.
Pure stage: output depends only on its input, and it changes nothing else. Same input, same output.
load, clean, aggregate: each stage a Parquet file in, a Parquet file out.
Your laptop goes to sleep while the clean stage is still writing tep_clean.parquet. You are left with half a file and no way to tell how much of it is good.
tep_clean.parquet
Idempotency: running a stage twice has the same effect as running it once, so you can safely retry a stage that failed.
If the clean stage is idempotent, the fix is to run it again. It reads the same input and overwrites the same output, so it makes no difference whether the first attempt stopped at row 100 or row 100,000.
In Lecture 3 the database's transactions did this for you.
One workload, one machine. Measure your own pipeline before you rewrite it.
l05-pipelines.ipynb
The Tennessee Eastman readings, one 4-stage pipeline built twice: pandas eager and Polars lazy, then timed.
The simulator's output has no gaps in it, so one clearly marked cell breaks the data on purpose: it deletes a chunk of readings and freezes one sensor at a constant. Now the cleaning stages have something real to fix.
pandas materializes a new table after every stage.
Polars builds a plan and runs it in one pass at .collect().
Print that plan with .explain(): the column selection is folded into the read. Lecture 4's DuckDB trick, no server.
.explain()
speaker: TEP is a Downs and Vogel 1993 benchmark, a real Eastman plant disguised. faultNumber 0 is normal, 1 to 20 are disturbances. Turning the raw log into per-fault means is a pipeline: load, clean, reshape, aggregate.