| known at |
not known at |
|---|---|
| the channel's own past | its future (that is the target) |
| valve positions now | valve positions later |
| a written setpoint schedule | where the operator will move it |
| tomorrow's weather forecast | tomorrow's weather |
Anything in the right column is leakage, stretched across
Autocorrelation at lag
statsmodels.graphics.tsaplots.plot_acf, or three lines of numpy
Pressure and separator level: run 1, Rieth et al. (2017). Random walk: simulated. figures/make_figures.py
| series | lag-1 ACF | best simple forecast |
|---|---|---|
| reactor pressure | 0.94, decays by ~75 min | recent past, then the mean |
| separator level | 0.02 | the mean: the controller removes every deviation |
| random walk (a price) | 0.98, stays high | the last value |
| its changes (returns) | -0.01 | nothing beats zero |
7 of 22 continuous TEP channels look like separator level: white noise.
Stationary: mean, spread and autocorrelation do not change over time. A plant at an operating point is roughly stationary; a random walk is not.
y[t] - y[t-1]A plant at an operating point has mostly neither.
A daily cycle there usually has a measured cause (cooling water temperature): put that column in the table, not the hour.
A stock price behaves like a random walk. On a held-out year, persistence beats your fitted model. What is the most likely explanation?
Persistence (the naive forecast): y[t+h] = y[t]. Lecture 7 used it without the name.
| baseline | forecast | where it is strong |
|---|---|---|
| persistence | the last value | short horizons, random walks |
| mean | the training average | long horizons, white noise |
| seasonal naive | the value one season ago | load, retail |
| drift | the average past change, extended | trending series |
For a stationary series with spread

Train runs 1 to 300, test runs 401 to 500. Test standard deviation 7.51 kPa. figures/make_figures.py
| horizon | persistence | mean | direct model | skill |
|---|---|---|---|---|
| 3 min | 1.92 | 7.51 | 1.83 | 5 % |
| 30 min | 5.82 | 7.57 | 5.02 | 14 % |
| 45 min | 7.39 | 7.59 | 5.87 | 21 % |
| 120 min | 12.29 | 7.65 | 7.17 | 6 % |
RMSE in kPa. The model earns most where neither free guess is good.
Skill score:
Hyndman and Koehler (2006), author's copy
Your one-step pressure model scores R-squared 0.99 on held-out runs. What should you compare it with first?
Exchange-rate models of the 1970s (money supply, interest rates, trade balances):
They fitted history well. Only the baseline comparison showed they could not forecast.
Meese and Rogoff, Fed IFDP 184 (1981), the working paper of the 1983 J. Int. Econ. article
100,000 real series, every team scored the same way. Benchmarks included Naïve2 and Comb (three simple exponential smoothers).
Makridakis, Spiliotis and Assimakopoulos (2018), IJF 34, abstract
| route | examples | strengths |
|---|---|---|
| statistics and control | AR, ARX, ARIMA, state space | small, well understood, standard errors |
| reduction to regression | lag table + ridge, forest, boosting | any regressor, many inputs, nonlinear |
statsmodels for the first route, scikit-learn and skforecast for the secondmodel = make_pipeline(StandardScaler(), Ridge(alpha=1.0))
model.fit(X_train, y_train) # scaler sees training rows only
Direct: one model per horizon, trained on y[t+h]. Recursive: one model for y[t+1], applied
u[t+1] ... u[t+h-1]; direct needs only what is known at t| horizon | direct | recursive | mean |
|---|---|---|---|
| 3 min | 1.83 | 1.83 | 7.51 |
| 30 min | 5.02 | 5.19 | 7.57 |
| 60 min | 6.51 | 7.08 | 7.61 |
| 120 min | 7.17 | 8.11 | 7.65 |
Test RMSE in kPa, runs 401 to 500. Same 10 lags, same ridge.
At two hours, recursive is worse than the mean.
Adding the 11 current valve positions to direct: 5.02 to 4.71 kPa at 30 min.
On the 111 series of the NN5 competition, recursive beat direct.
Taieb, Bontempi, Atiya and Sorjamaa (2012), arXiv preprint
Lecture 7 showed rolling_mean. The open question was the window length.
| window | follows | costs |
|---|---|---|
| short | the latest move | carries the noise |
| long | the level | lags behind a change |
The horizon is the criterion: try two or three lengths, keep what lowers the held-out error at your
The score measures how densely you sampled, not how well you forecast. Lecture 7 named this train-test contamination.

One run at a time, figures/make_figures.py
The boundary: pooled over 200 independent runs, ridge scores 4.74 shuffled and 4.71 by run. K-fold is valid for a purely autoregressive model with uncorrelated errors.
Shuffling bites hardest on one short series, a flexible model, correlated errors.
Bergmeir, Hyndman and Koo (2018), author's copy
| bias | in a backtest | in a plant |
|---|---|---|
| look-ahead | using data not available on the decision date | a shuffled split, a future valve |
| survivorship | testing on today's index members only | keeping only the runs where nothing went wrong |
Both inflate the score. Neither raises an error.
Anomaly detection: flag observations that do not fit normal data. A residual detector alarms when the forecast error leaves its normal range.
Pressure, one step: residual sd 1.70 kPa, threshold 4.38 kPa. No fault examples needed.
False-alarm rate: how often it fires on normal data. Detection delay: time from fault onset to first alarm. Lowering one raises the other.

Fault 1 (A/C feed ratio step), run 1, faulty training file of Rieth et al. (2017). 3-in-a-row alarm 45 min after onset.
One channel sees only what that channel sees.
| failure | what happens | what helps |
|---|---|---|
| regime change | new grade, catalyst, market | watch residuals, retrain |
| feedback | controller (Lecture 7), traders acting on forecasts | log setpoints, retrain after retuning |
| unseen faults | extrapolation from a regime never trained on | detect, do not forecast through |
| horizon ceiling | pressure skill 6 % at two hours | the ACF shows it before fitting |
Each one returns a confident number, not an error.
l08-forecasting.ipynbReactor pressure, 500 runs. Build the horizon table, score persistence and the mean,
fit direct and recursive, then shuffle and don't.
Nicknames only. Everyone who skipped one still counted in every bar you saw.
Assignment 4 is released today, due 2026-09-28: an h-step forecaster for stripper temperature
Practice module for this session, for participation credit
Reading FPP3 chapter 5, https://otexts.com/fpp3/toolbox.html
Full notes, with all sources: lectures/l08/notes.md
Budget: 90 minutes of lecture, then 20 for the notebook and questions. Lecture: opening 6, supervised 9, forecastable 14, baselines 17, models 13, evaluation 15, residuals 8, limitations 5. About 87, with recap and standings after the notebook. Abort order if running long: window-features slide, look-ahead/survivorship slide, the second residual slide.
Roadmap, 1 minute. Say where the demo sits: the lecture runs to the recap, then the last 20 minutes are the notebook and questions.
The hook. A one-step model on pressure reports R-squared near 0.99 and is mostly the free guess. Ask who would sign off on that model. Do not give numbers yet; slide 21 has them.
One minute. Ask the room for the free guess in their own field before reading the table. The finance row returns in the clicker on slide 16.
About 9 minutes for this arc. No new mathematics: L7 already solved a least-squares problem.
Point back at L7's lstsq. Same step, better bookkeeping. fit/predict is the only new vocabulary, and the one-line swap between model families is what makes the rest of the course possible.
The consequence is the part students miss: a dataset used to CHOOSE a setting can no longer measure it. A4 makes them do this for real, nine settings scored on runs 301 to 400.
Write y[t+h] on the board. The only code change from L7 is shift(-h). Ask what .over(RUN) protects against, and let someone say the run boundary.
The weather row is the one that lands: you may use tomorrow's FORECAST, not tomorrow's weather. Ask for the plant equivalent; a written setpoint schedule is the answer.
About 14 minutes. This section decides whether a model is possible at all, before any fitting.
Ask them to sketch the ACF of a constant series and of pure noise before the next slide. Thirty seconds, and it makes the next figure readable.
speaker: ask the room to say, for each column, what they would guess for the next sample.
Ask the third row before revealing it: the random walk's ACF is 0.98, so is it the easiest to forecast? The returns line answers it. Same trap as the R-squared of 0.99 in the opening.
Differencing is the fix, and prices to returns is the same move. ARIMA is named here, not taught; say so, so nobody waits for it.
A continuous plant has no weekly seasonality. Its daily cycle is usually cooling water temperature, so put the measured channel in the table rather than the hour of the day.
Tests whether a random walk is read as "persistence is near optimal". A tempts students who assume a loss to a baseline means broken code; D tempts those who think the mean should be used, which is useless on a drifting level.
About 17 minutes. The core of the session: everything here is what makes a score mean something.
L7 used persistence without naming it. Ask which baseline they would pick for a stock price, a load curve, and a controlled level, and let the disagreement stand until the next slide.
speaker: the ACF you just looked at tells you the crossover before any fitting.
Walk the four curves, do not just show them. Persistence rises 1.92 to 12.29 kPa; the mean is flat at about 7.5; they cross between 45 and 60 minutes. The model sits under both and is closest to them at the two ends. Recursive passes the mean past about 75 minutes, which sets up the next section.
Read the 5 % row and the 21 % row aloud. The model earns least where a free guess is already good, and most where neither one is.
The reference is the BETTER baseline at that horizon. MASE if asked: scale-free, below 1 beats the naive forecast. Always say the kPa as well; an operator can judge 5 kPa, not 0.14.
D is a real baseline but the weak one at 3 minutes (7.51 against 1.92), which is the point.
The quote is the slide. Note the rolling regressions: rolling-origin evaluation, in 1983. Ask why beating a random walk on exchange rates is so hard, and link back to slide 13.
speaker, the takeaway for both case studies: report persistence and the mean at every horizon you claim, beat the better one or say you did not, and if you did not, ship the baseline. It is cheap, cannot overfit, and needs no maintenance.
About 13 minutes.
Two routes that meet at the linear model: ridge on ten lags IS an AR(10). ARIMA is named only.
Contrast with L7: plain least squares does not care about scale, ridge does, because the penalty is in the coefficients' units. The pipeline makes the rule structural rather than remembered.
Ask how to reach 10 steps with a one-step model. Someone will say iterate. That is the recursive strategy, and the question of what it is standing on at step 7 follows by itself. The input line needs saying out loud, because the bullet is compressed. Write the two steps up: y[t+1] = f(y[t], y[t-1], ..., u[t]) y[t+2] = f(y[t+1], y[t], ..., u[t+1]) <- u[t+1] does not exist yet You have the prediction for y[t+1]; you do not have the valve positions three minutes from now, because the controller has not moved them. Four ways out, and we take the last: forecast the valves too (a second model, errors compounding twice, and the controller reacts to exactly what you are forecasting); freeze them at u[t] (assumes the controller does nothing for two hours, and it is no longer the model you fitted); use a plan, which is legitimate when a setpoint schedule or a recipe exists, the same exception as tomorrow's weather forecast; or drop the inputs, which is what our recursive model does, lags of pressure only. That is why direct can take the valves and gain something real: 5.02 kPa to 4.71 at 30 minutes. A4 section 4 asks students for this argument.
Say the units: every number in the table is a test RMSE in kPa, on the held-out runs, against a channel whose own spread is 7.51 kPa. At two hours the recursive forecast (8.11) is worse than simply predicting the mean (7.65). Direct costs one model per horizon, which for a linear model is nothing.
Hold this line: our data says direct, the NN5 competition says recursive. The transferable rule is to measure both, not to memorize a winner.
speaker: first slide to cut if running long
About 15 minutes, and the notebook runs it again afterwards.
Ask where a test row's neighbours end up after shuffling. Three minutes either side, in the training set. The ACF from slide 12 is why that matters.
Let the room read the bars before you say anything, then give them the reading rule: these are errors, so lower is better, and the dashed line at 6.09 kPa is what persistence costs on the same folds. Below the line means the model is worth having; above it means the free forecast was better. Then the question to put to them: which of these two pictures would you show a plant manager? The left pair is the same model as the right pair. Nothing changed but how the rows were split. Numbers are on the next slide.
The honest split says do not ship this model. Then give the boundary immediately: pooled over 200 runs, ridge scores 4.74 shuffled against 4.71 by run. Nobody should leave thinking a shuffled split is always fatal; it bites on one short series with a flexible model.
speaker: second slide to cut if running long
speaker, say this over the divider: a residual is measurement minus prediction. A good forecaster leaves white noise, so plot the residual ACF. A spike at one lag means a missing feature at that lag; a slow decay means a missing input. The notes carry the details.
The threshold comes from held-out fault-free runs, never from the training residuals. Residual sd 1.70 kPa, 99th percentile 4.38. No fault examples are needed anywhere in this recipe.
Do the arithmetic out loud: 1 % of 480 samples a day is 4.8 false alarms a day, and an operator will mute that within a week. Three in a row drops it to 0.01 a day, for at least six minutes of delay.
Fault 1 is a step in the A/C feed ratio, one hour into the run. The three-in-a-row detector fires 45 minutes after onset. Let them look before you explain the residual returning to the band.
speaker: third cut if running long
Every row of this table returns a confident number and no error message. The feedback row is L7's closed-loop case arriving again.
The last 20 minutes, notebook then questions. Pause at "stop and predict" before the baseline table prints. The last cell takes about 15 seconds.
Land these six lines fast, then go to the notebook. The last 20 minutes are demo plus questions.
Skip this slide if no clicker questions were run.
A4 is released today and due 09-28. The data is a 25 MB Parquet from kitchin-services; the link is in the assignment, and the checksum check is one line.