Your data is discrete. One row every 3 minutes, nothing in between. A dataframe has rows, not a continuum.
So a model you can fit to a table has to be in discrete time.
And the discrete model is what tells you which columns the table needs.
06-763 / L7
Continuous vs discrete? the continuous model
A vessel or line with one place to store something (heat, mass, momentum):
Time constant: how long to cover 63 % of the distance to the new steady state. Steady-state gain: how far it finally moves per unit of input.
06-763 / L7
Continuous vs discrete?, discretizing it
We need sample from sample . That needs an assumption about the input in between, and the plant gives it to us free: the valve holds its position between samples.
That is zero-order hold (ZOH). Integrate across one interval with u held constant:
06-763 / L7
Continuous and discrete, the model gives the feature list
Read that equation as a table specification, not as physics.
To predict row you need exactly two numbers from row :
the model asks for
the column
the previous measurement, xmeas_4_prev
the valve at that row, xmv_4
That is where the features come from. Not a guess, not everything in the file. The discretized model names them.
06-763 / L7
Continuous and discrete, terminology (ARX)
ARX(1,1): autoregressive with exogenous input. One lag of the output, one of the input. The simplest model structure in system identification.
06-763 / L7
Continuous and discrete, the lag plot
A straight line whose slope is a, simulated with the valve held steady. Read the coefficient off the chart before you fit anything.
06-763 / L7
Continuous and discrete, getting the physics back
Time constant: how long the loop takes to cover about 63 % of the distance to its new steady state.
Steady-state gain: how far the output finally moves per unit of input.
06-763 / L7
Leakage
06-763 / L7
Leakage, the definition
Data leakage: a model is trained on information that would not be available at the time it has to make a prediction.
faultNumber: what was wrong during the run, written down afterwards
not available at prediction time
in production it is the thing you are predicting
And the subtle one: xmeas_23 to xmeas_41, the slow analysers.
06-763 / L7
Leakage, two ways it gets in
Target leakage. A column holds information from after the row's timestamp. The chargeback, faultNumber, the analysers. A property of the table. Ours today.
Train-test contamination. Training rows come from after the test rows. Shuffle a time series at random and every test row has its neighbours in training.
The second needs a held-out score to show, so it belongs with the sessions that fit models. Carry the fix: split on time, not at random.
06-763 / L7
Common data preparation operations
06-763 / L7
Data preparation, build the clock and check it
No timestamps in this file. Just sample, an integer counter. Build the clock, then check it:
Why bother:shift(1) reaches back one row, not three minutes. If one row is missing, that lag quietly spans six minutes instead, and the fitted time constant is wrong.
06-763 / L7
Data preparation, putting rows back on the grid
Rows are missing. Put them back, so every row is 3 minutes apart:
build the regressor matrix, twice, at two lag depths
solve with lstsq, convert to a time constant
freeze the valve and refit
06-763 / L7
What to watch
The matrix, printed. Read one row across: these were the flow and the valve at this moment, and that is what the flow did next.
It loses rows. 480 in; one lag leaves 479, three lags leave 477. A row needs all of its past and its answer.
The coefficient splits. With one lag, y[t] gets 0.753. With three, it gets 0.417 and the rest lands on y[t-1] and y[t-2].
The gain goes to zero when the valve is frozen.
06-763 / L7
Fitting, and converting to physics
06-763 / L7
Fitting, deviation variables
y = y - y.mean() # xmeas_4 sits at 8.79
u = u - u.mean() # xmv_4 sits at 57.6
Both columns are almost entirely offset. A model with no constant term spends both coefficients explaining that offset and gets a badly wrong.
In process control this is working in deviation variables.
06-763 / L7
Fitting, what came back
fitted
a
0.7529
b
0.0322
tau
10.6 min
K
0.1304 kscmh per % of valve
Two coefficients in, minutes and flow-per-percent out.
Minutes, and flow per percent of valve. A plant engineer understands this.
06-763 / L7
Multiple lags and the regression vector
06-763 / L7
Multiple lags, the general form
One lag is where you start, not a rule. There are situations you may need more lag/regressors.
df.with_columns([
pl.col("xmeas_4").shift(k).over(GROUP).alias(f"xmeas_4_lag{k}")
for k in (1, 2, 3)
])
Nothing about the method changes.
06-763 / L7
Multiple lags, the regression vector
Collect the right-hand side into one row, the regression vector:
Stack one row per sample and you have the design matrix. The whole model is
lstsq solves that for at any width of .
06-763 / L7
Model structures: ARX, NARX, state space, RNN
06-763 / L7
Model structures, the family
ARX(1,1)
what you just built
ARX(, )
more lags, same lstsq
FOPDT
add dead time; what PID rules are written against
NARX
same , nonlinear function on it
state-space
; ours is the 1x1 case
RNN / LSTM
vector state, nonlinear update, gates
06-763 / L7
Model structures, the same arithmetic
Ours:
An RNN cell:
Some of where you were, plus some of the input. An LSTM adds gates deciding how much of to keep, which is why its name says memory.
What changes across the family is how much state you carry and how nonlinear the update is.
06-763 / L7
Model structures, why the table came first
Every model in that list uses the same feature table you just built.
If the past is missing from the table, or a column leaked, no amount of capacity further down the list repairs it.
A bigger model cannot recover information the table never contained.
06-763 / L7
Limitations: what the data cannot answer
06-763 / L7
Limitations, the valve that never moved
Same code, same clean table, one different day: the operator left the valve alone all day.
rank of the design matrix: 2 -> 1
b = 0.0000 K = 0.000 (the truth is K = 0.130)
The data really does contain no evidence about the valve. It never moved, so the data cannot say what it would have done.
06-763 / L7
Limitations, persistent excitation
Persistent excitation: the input has to move over enough distinct frequencies to make the parameters identifiable. A constant input is persistently exciting of order zero.
How much is enough? A single step is plenty. What matters is the rank: a constant input gives rank 1 and no gain, a single step gives rank 2 and a usable fit.
The fix is an experiment, not an algorithm. Move the setpoint, or add a small deliberate wiggle (a PRBS).
The controller moves the valve because the measurement moved. In the record, the effect arrives first and the cause follows.
This is closed-loop identification, and it is the normal condition of historian data.
06-763 / L7
Limitations, what to do
log the setpoint as well as the valve, so the thing that moved on its own is in the table
fit on days where the setpoint stepped, not where the plant held still
ask for a deliberate wiggle: minutes of a small perturbation buys a number no logged data contains
06-763 / L7
Recap
A dynamic loop draws a loop, so its past has to become a column
Leakage: every value in a row was knowable at that row's timestamp
The group key travels with the shift, and here it is two columns, not one
Two coefficients convert to a time constant and a gain: 10.7 min on our record, 10.6 on the plant's
A valve that never moved, and a controller in the loop, both return confident wrong numbers
06-763 / L7
Recap, the names you can now look up
You built an ARX(1,1) model by zero-order hold discretization, widened it to a regression vector, and met two reasons the data refuses to answer.
None of this is new. Process control has been identifying models from plant data for over fifty years, and the ARX structure you just built is where that field starts. What changed is the size and "nature" of the models, not the table underneath them.
Nelles, Nonlinear System Identification: From Classical Approaches to Neural Networks, Fuzzy Models, and Gaussian Processes, 2nd ed., Springer 2020. Goes from simple ARX to neural networks.
06-763 / L7
Next
Practice module for this session, for participation credit
Assignment 4 is released at the next session, on 2026-09-21
speaker: this is the slide the whole session hangs on. Ask the room how they would catch
any of those three from a score alone. Let the silence sit.
speaker: point at FC-4 on the Feed A,B,C line. That single loop is the whole session. Then
point at the three Analyzer blocks on the edges, because those come back in the channel screen.
speaker: ask the room who has actually been told no. Then wait for it.
TIMING, 110 minutes. 60 slides, of which 8 are dividers.
opening through "the answer is a column" ~14 min
one loop / continuous vs discrete ~22 min
leakage ~10 min
data preparation (rapid fire) ~ 8 min
demo ~35 min
fitting, more lags, model structures ~15 min
what the data cannot answer ~ 8 min
recap ~ 4 min
NO clicker questions now. All three were cut, and the leaderboard slide went with them.
clicker-slide.js is still loaded at the end of the deck, deliberately: CI copies it beside
every deck anyway, and leaving the tag means a clicker slide pasted back in just works.
ABORT SEQUENCE, in this order of preference:
1. "Data preparation, rolling window statistics", replaced by its one sentence
2. "Model structures, the same arithmetic" (the family table carries the bridge alone)
3. "Data preparation, what the plateaus are" (the screen slide before it carries the number)
NEVER cut the closed-loop block. The logged-data cloud early on is its setup.
NOTE: the run-boundary and group-key slides were cut from the deck on 2026-09-16. That
content (the .over("simulationRun") half-key bug, and the null-count check) now lives only
in the notes and in the demo, where it is run live. Do not let it drop out of the demo.
speaker: do NOT explain this yet. Say you will come back to it and move on. It is the setup for
the closed-loop block at the end. There is no callback slide any more, so when you reach
"Limitations, closed-loop identification", say out loud that this cloud was the evidence.
speaker: the cheapest diagnostic in the course. Do not cut. A curved lag plot says the model is
wrong; a fat cloud says the signal is mostly noise.