Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Validation

This is the part of the project that determines whether anyone trusts the result, so it gets the most care. The strategy is a ladder: prove the pieces exactly, prove the derivative to near machine precision, prove the stochastic call order exactly, then prove the long-run behaviour statistically and, finally, prove it on the downstream task people actually care about.

Every number on this page was measured, and each one names the LOG.org iteration that measured it. None of them is a target, a plan, or a rounded recollection. The project's operating rule is to record numbers rather than verdicts, precisely so that a degradation inside tolerance is still visible: if a maximum relative error moves from 3e-14 to 8e-13, both pass a 1e-12 gate and something has broken, and the only place that is visible is the logged history.

The figures are held to the same rule. Each one is drawn by cargo xtask validate from the run's own output, committed alongside the generated chapters, and captioned with the condition that would make it false. Nothing in one is placed by hand: a marker is orange because its value is on the wrong side of a gate, never because a test's name was recognised, so a genuine regression and a deliberate positive control are drawn identically and the caption is what tells them apart.

All numbers below were produced with gfortran 15.2.0 and the pinned rustc 1.97.1. The oracle's compiler flags are fixed in build.rs and asserted by a test; changing them invalidates every Tier 1 and Tier 2 number here, so it is a logged re-baseline rather than a casual edit.

The oracle harness

tepsim-oracle is a development-only crate whose build.rs compiles the unmodified teprob.f and temain_mod.f with gfortran and links them through a small C shim exposing TEINIT, TEFUNC, each TESUBn, and read and write access to every COMMON block. A Rust test can therefore set the Fortran into an arbitrary state, call it, and compare against the Rust implementation in the same process.

That crate is never a dependency of tepsim, tepsim-py or tepsim-wasm. It runs on Linux and macOS runners where gfortran is available, and building it first is what converts the whole port from "read carefully and hope" into a differential testing exercise.

Alongside it sits a cheaper check that needs no Fortran at all. cargo xtask fidelity runs the port forward 100 steps from the nominal state and diffs states, derivatives, measurements and the generator word against a golden trace committed to the repository. It takes about a second and runs at the top of every session. Since B-0026 it has reported 100 of 100 steps diffed, worst 3.521e-14 at YP(12) step 77 against a 1e-12 gate, and that same number has now been recorded unchanged for eleven consecutive iterations, most recently in the entry for B-0052.

Tier 1: the utility routines, exactly

TESUB1 (enthalpy), TESUB2 (temperature from enthalpy by Newton), TESUB3 (heat capacity) and TESUB4 (liquid density) are swept over a simplex grid, ten million random Dirichlet samples, and a boundary pool, at every temperature in the physical range, for each of the three ITY modes. The gate PLAN.org sets is a maximum relative error below 1e-13 with a ULP histogram reported rather than a pass or fail.

The measured result is not "inside 1e-13". It is zero.

routineITYcasesmax relative errormax ULPhistogramfrom
TESUB109,987,4900.000e000:9987490B-0009
TESUB119,987,4900.000e000:9987490B-0009
TESUB129,987,4900.000e000:9987490B-0009
TESUB309,987,4900.000e000:9987490B-0009
TESUB319,987,4900.000e000:9987490B-0009
TESUB329,987,4900.000e000:9987490B-0009
TESUB4n/a9,987,4900.000e000:9987490B-0010

That is 59,924,940 evaluations across TESUB1 and TESUB3 with zero differing bits, p50 = p90 = p99 = p100 = 0 in all six sweeps. A separate test asserts bit equality directly, because a 1e-13 threshold would let a drift to 1e-15 pass unnoticed forever.

The 1e-13 gate is therefore not what is holding the line, and the entry for B-0009 says why that is the right expectation: both routines are straight-line arithmetic over constants already proved bit-identical, so once the association and the literal precisions are right there is nothing left to differ by. Any future Tier 1 routine that lands at 1e-15 rather than at zero has something wrong with it that a tolerance would hide.

The sabotage check. Substituting a double-precision 273.15 for the widened f32 at teprob.f:1411, and changing nothing else, gives a maximum relative error of 1.597e-5 at grid#704 T=21.875, 103,098,852,352 ULP (B-0009). That is the size of the error a single mis-transcribed literal produces, and it is why constants in this project are transcribed and asserted rather than retyped.

TESUB2 is the Newton solve, and it is measured from two starting strategies (B-0011). A warm start, which is what every call site actually does, converges in one step and recovers the answer exactly: round-trip error 0e0. A cold start from the far end of the range is what exercises the iteration, and its worst round trip is 5.12e-13 C, at dirichlet#720561 T=171.86, face#1309309 T=146.56 and face#96 T=144.14. Over 59,924,940 solves there were zero differing bits and zero abandoned iterations, which is the measurement that makes delta D-001's effect zero on the physical domain.

TESUB7, the random generator, must be exact and is (B-0005): 10^7 draws from the compiled-in seed with the XOR fold and the final state both matching; 10^6 draws each from five dataset seeds; draw-by-draw comparison over 200,000 draws for four seeds with every draw and every intermediate state bit-identical; and interleaved output modes over 50,000 draws on an irregular pattern. Its exactness rests on reproducing the rounding rather than removing it: the product exceeds 2^53 on 0.7716 of draws, against 0.7728 predicted from the arithmetic, so a "fixed" integer recurrence diverges at draw 0.

TESUB5 and TESUB6 are exact in both libm configurations, since neither touches a transcendental, and their draw counts were recovered independently by stepping a port-side generator from the word before each call to the word after: 3 draws for TESUB5, 12 for TESUB6, on every case and for both flag values (B-0028).

Tier 2: single-step derivative equivalence

Both implementations are forced into an identical state, all fifty states, the twelve manipulated variables, the twenty IDV flags, the full walk state, the generator word and the nineteen held analyser readings, evaluated once, and compared on all fifty components. Sampling is from three pools: states along the nominal closed-loop trajectory, random perturbations of those states scaled across several orders of magnitude, and adversarial states placed deliberately at every discontinuity and clamp in the model.

The tolerance is relative to the scale of the terms, not to the result, and that is a decision with a measurement behind it rather than a relaxation. A balance is inflow minus outflow, and near steady state those nearly agree. YP(2), the inert's reactor balance, is a difference of two flows around 660 whose result is a few parts in ten thousand of either. One ULP of difference in each term, which is all the vendored libm costs, is 1e-16 of the terms and 1e-4 of the result. Measured against the result, 28 of the 50 components exceed 1e-12 while the whole right-hand side is bit-identical to gfortran under libm-system. Each balance therefore reports the magnitude of its largest term alongside its value, and the gate is the error over that.

Acceptance, from B-0026: all fifty components, 2,412 running states, all three pools, worst 6.093e-14 at YP(7), perturbed#300, against a 1e-12 gate.

The whole tier fits in one picture. Every comparison it makes, run twice, once against each libm, plotted at its own maximum relative error:

GENERATED by `cargo xtask validate --tiers 1,2,3 --smoke` from commit `b6ef4ee-dirty`. Do not edit by hand: the next run overwrites it. Tier 2: every comparison against the 1e-12 gateA logarithmic strip plot. Each dot is one comparison against the Fortran, placed at its maximum relative error, in a lane named for the test target that produced it. 115 of 141 comparisons are exactly zero and sit in the separate lane at the left. The gate at 1e-12 is a dashed vertical line with the region beyond it shaded; 1 dots lie beyond it. Tier 2: every comparison, against the gate 141 comparisons over 9 target(s). 115 are bit-identical to the Fortran. 1 lies beyond the gate. 1e-16 1e-14 1e-12 1e-10 1e-8 1e-6 1e-4 1e-2 maximum relative error against the Fortran = 0 gate 1e-12 tier2_unpack tier1 TCV with the carried seed: exactly 0 (vendored libm, solving_from_a_fixed_guess_instead_of_the_seed_breaks_bit_equality) tier1 TCV from a fixed guess: 9.901e-16 (vendored libm, solving_from_a_fixed_guess_instead_of_the_seed_breaks_bit_equality) tier1 unpack UCLR: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack UCLS: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack UCLC: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack UCVV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack UTLR/UTLS/UTLC/UTVV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack XLR: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack XLS: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack XLC: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack XVV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack ESR/ESS/ESC/ESV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack TCR/TCS/TCC/TCV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack TKR/TKS/TKV: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack DLR/DLS/DLC: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier1 unpack VLR/VLS/VLC: exactly 0 (vendored libm, the_unpacking_matches_the_fortran_over_all_three_pools) tier2_equilibrium tier1 equilibrium VVR/VVS: exactly 0 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PPR: 2.802e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PPS: 2.784e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PTR/PTS/PTV: 3.578e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium XVR: 4.843e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium XVS: 4.302e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UTVR/UTVS: 4.443e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UCVR: 5.077e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UCVS: 6.230e-16 (vendored libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 PPR/PPS for A, B and C: exactly 0 (vendored libm, the_ideal_gas_partial_pressures_are_bit_identical_under_the_vendored_libm) tier1 PTV: exactly 0 (vendored libm, the_ideal_gas_partial_pressures_are_bit_identical_under_the_vendored_libm) tier1 equilibrium VVR/VVS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium PPR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium PPS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium PTR/PTS/PTV: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium XVR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium XVS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium UTVR/UTVS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium UCVR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium UCVS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_agrees) tier1 equilibrium VVR/VVS: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PPR: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PPS: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium PTR/PTS/PTV: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium XVR: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium XVS: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UTVR/UTVS: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UCVR: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier1 equilibrium UCVS: exactly 0 (platform libm, the_equilibrium_matches_the_fortran_over_all_three_pools) tier2_kinetics tier1 CRXR(2), never assigned: exactly 0 (vendored libm, the_inert_has_no_net_production_in_either_implementation) tier1 kinetics RR: 8.492e-16 (vendored libm, the_kinetics_match_the_fortran_over_all_three_pools) tier1 kinetics CRXR: 8.492e-16 (vendored libm, the_kinetics_match_the_fortran_over_all_three_pools) tier1 kinetics RH: 7.419e-16 (vendored libm, the_kinetics_match_the_fortran_over_all_three_pools) tier1 kinetics RR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 kinetics CRXR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 kinetics RH: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 CRXR(2), never assigned: exactly 0 (platform libm, the_inert_has_no_net_production_in_either_implementation) tier1 kinetics RR: exactly 0 (platform libm, the_kinetics_match_the_fortran_over_all_three_pools) tier1 kinetics CRXR: exactly 0 (platform libm, the_kinetics_match_the_fortran_over_all_three_pools) tier1 kinetics RH: exactly 0 (platform libm, the_kinetics_match_the_fortran_over_all_three_pools) tier2_streams tier1 streams XST (10 streams x 8): 4.736e-16 (vendored libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams XMWS (the 6 that exist): 4.370e-16 (vendored libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams TST: exactly 0 (vendored libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams HST: 5.050e-16 (vendored libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams XST (10 streams x 8): exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 streams XMWS (the 6 that exist): exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 streams TST: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 streams HST: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 streams XST (10 streams x 8): exactly 0 (platform libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams XMWS (the 6 that exist): exactly 0 (platform libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams TST: exactly 0 (platform libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier1 streams HST: exactly 0 (platform libm, the_stream_table_matches_the_fortran_over_all_three_pools) tier2_flows tier1 flows FTM (the 10 assembled streams): 2.246e-14 (vendored libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows FCM (10 streams x 8): 2.268e-14 (vendored libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows FWR/FWS/AGSP: exactly 0 (vendored libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows CPDH: 2.034e-15 (vendored libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows HST(9) after the compressor bump: 6.522e-15 (vendored libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 UAC, via QUC: exactly 0 (vendored libm, the_steam_coefficient_matches_through_the_condenser_duty) tier1 flows FTM (the 10 assembled streams): exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 flows FCM (10 streams x 8): exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 flows FWR/FWS/AGSP: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 flows CPDH: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 flows HST(9) after the compressor bump: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 flows FTM (the 10 assembled streams): exactly 0 (platform libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows FCM (10 streams x 8): exactly 0 (platform libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows FWR/FWS/AGSP: exactly 0 (platform libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows CPDH: exactly 0 (platform libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 flows HST(9) after the compressor bump: exactly 0 (platform libm, the_flow_network_matches_the_fortran_over_all_three_pools) tier1 UAC, via QUC: exactly 0 (platform libm, the_steam_coefficient_matches_through_the_condenser_duty) tier2_stripper tier1 stripper SFR: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FTM (5, 12): the column's own outlets: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FTM (7): the reactor-inlet alias: 1.484e-15 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FCM (5, 12): the column's own outlets: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FCM (7): the reactor-inlet alias: 1.610e-15 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper XST (5, 12): the column's own outlets: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper XST (7): the reactor-inlet alias: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper TST (5, 7, 12): exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper HST (5, 12): the column's own outlets: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper HST (7): the reactor-inlet alias: exactly 0 (vendored libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper SFR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper FTM (5, 12): the column's own outlets: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper FTM (7): the reactor-inlet alias: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper FCM (5, 12): the column's own outlets: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper FCM (7): the reactor-inlet alias: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper XST (5, 12): the column's own outlets: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper XST (7): the reactor-inlet alias: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper TST (5, 7, 12): exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper HST (5, 12): the column's own outlets: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper HST (7): the reactor-inlet alias: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 stripper SFR: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FTM (5, 12): the column's own outlets: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FTM (7): the reactor-inlet alias: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FCM (5, 12): the column's own outlets: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper FCM (7): the reactor-inlet alias: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper XST (5, 12): the column's own outlets: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper XST (7): the reactor-inlet alias: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper TST (5, 7, 12): exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper HST (5, 12): the column's own outlets: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier1 stripper HST (7): the reactor-inlet alias: exactly 0 (platform libm, the_stripper_matches_the_fortran_over_all_three_pools) tier2_heat tier1 heat UAR: exactly 0 (vendored libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUR: exactly 0 (vendored libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUS: 9.572e-13 (vendored libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUC: exactly 0 (vendored libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat UAR: exactly 0 (platform libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUR: exactly 0 (platform libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUS: exactly 0 (platform libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat QUC: exactly 0 (platform libm, heat_transfer_matches_the_fortran_over_all_three_pools) tier1 heat UAR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 heat QUR: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 heat QUS: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 heat QUC: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier2_measurements tier1 XMEAS(1..22), noise-free: 8.774e-15 (vendored libm, the_measurements_match_the_fortran_over_all_three_pools) tier1 ISD as 0 or 1: exactly 0 (vendored libm, the_shutdown_detector_agrees_with_the_fortran_on_every_state) tier1 XMEAS(1..22), noise-free: exactly 0 (platform libm, the_algebra_is_bit_identical_once_exp_and_pow_agree) tier1 XMEAS(1..22), noise-free: exactly 0 (platform libm, the_measurements_match_the_fortran_over_all_three_pools) tier1 ISD as 0 or 1: exactly 0 (platform libm, the_shutdown_detector_agrees_with_the_fortran_on_every_state) tier2_balances tier1 YP(1..50), relative to the derivative (reported): 1.393e-4 (vendored libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), relative to the scale of the terms (the gate): 6.093e-14 (vendored libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), plant frozen: exactly 0 (vendored libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), relative to the scale of the terms (the gate): 6.093e-14 (vendored libm, tier2_acceptance_table) tier1 YP(1..50), relative to the derivative (reported): exactly 0 (platform libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), relative to the scale of the terms (the gate): exactly 0 (platform libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), plant frozen: exactly 0 (platform libm, all_fifty_derivatives_match_the_fortran_over_all_three_pools) tier1 YP(1..50), relative to the derivative (reported): exactly 0 (platform libm, the_whole_right_hand_side_is_bit_identical_once_exp_and_pow_agree) tier1 YP(1..50), relative to the scale of the terms (the gate): exactly 0 (platform libm, tier2_acceptance_table) vendored libm vendored libm platform libm platform libm, bit-exact claim beyond the gate beyond the gate hover a dot for the comparison it came from
Every Tier 2 comparison, against the 1e-12 gate. One lane per test target, one dot per comparison, on a logarithmic axis with a separate lane at the left for the comparisons that are exactly bit-identical. Blue dots are the libm-system build, where both sides call the same exp and the claim is bit equality rather than a tolerance; every one of them is in the zero lane, and one leaving it would falsify that claim. Grey dots are the vendored libm, and every one of them is inside the gate, though one in tier2_heat sits close enough to it to be worth hovering over: a table of maxima hides which comparison is nearest the line, and this does not. Any dot crossing into the shaded region would be a failure. The figure is written by cargo xtask validate whenever it runs Tier 2, from that run's own output; the measurements are the ones recorded for B-0026 in LOG.org, and the generated Tier 2 chapter carries the same figure with the exact command and commit that drew it, plus the same data as a table.
YPerror / scaleerror / valueratio
23.811e-141.393e-43.7e9
304.552e-152.487e-65.5e8
388.696e-151.187e-61.4e8
143.113e-142.618e-68.4e7
76.093e-147.835e-71.3e7
19-27, 37, 39-500.000e00.000e01x

28 of 50 components cancel by more than 100x, and the 22 that do not are exactly 0.000e0: bit-identical rather than merely inside the gate. The acceptance test asserts that contrast and fails if the count of heavily cancelling components drops below twenty, so the reasoning behind the decision expires loudly if the model ever stops cancelling.

Tier 3: RNG call-order equivalence

Both sides are instrumented to emit every draw, and the traces are diffed. This test exists specifically because it is the one the existing Python port would fail: its documented divergence is attributed to "the exact sequence of calls differs due to implementation details", which is exactly the defect a trace diff catches on the first run instead of after a 48-hour statistical comparison.

From B-0029, draw counts at four points in a run, with exact agreement at every one:

timedrawswhat is drawing
00noise skipped, walks reset
1e-6264noise only
0.15462noise, walk advance, gas analysers
0.30522and the product analyser

Scalings in one real evaluation: 30 signed and 432 unit. The 30 is 9 x 3 + 3, the walk advance; the 432 is 36 compositions at twelve draws each. Worst evaluation is 462 draws against a trace buffer capacity of 4096. The instrument is checked against the generator word rather than only against the values, which covers completeness, ordering and fidelity at once, and it holds over 113,088 draws.

Two independent methods agree on the counts. B-0027 measured them with no instrumentation at all, by stepping a port-side generator from the word before a call to the word after, and the trace lengths match that census exactly. Either method alone would be a claim; the two agreeing is evidence.

Tier 4: trajectory equivalence, diagnostic

Tier 4 is a diagnostic, not a gate. Long-horizon divergence between two programs that use different exp implementations is expected. PLAN.org asks for two things: that the error stay below the corresponding measurement noise standard deviation XNS(i) for at least the first several hours, and that the onset of divergence be explained by showing that switching libm moves it.

Open loop, nominal, 8 simulated hours (28,800 steps), from B-0034:

libmworst error, as a fraction of XNS(i)ever outside XNS
vendored5.149e-5never
platform0.000e0never

All 21 scenarios at 4 hours each stay within XNS for the whole run in both configurations, and every one is exactly 0.00e0 on the platform libm. Worst error by scenario on the vendored libm at 4 hours: nominal 2.45e-5, IDV(10) 3.66e-8, IDV(8) 1.42e-8, IDV(13) 1.05e-8, and the rest below 1e-8.

GENERATED by `cargo xtask validate --tiers 4` from commit `b6ef4ee-dirty`. Do not edit by hand: the next run overwrites it. Tier 4: trajectory error against the instrument noiseA logarithmic strip plot with one row per scenario. Each marker is the worst disagreement between the port and the Fortran over a 4 hour run, divided by the measurement noise standard deviation of the channel it occurred on. The band at one, where the error would equal the instrument noise, is shaded; 0 of 42 markers reach it. 21 are exactly zero. Tier 4: how far apart the trajectories get, in units of instrument noise 21 scenarios, 4 h each. Worst error over the run divided by XNS(i). 21 of 42 runs are bit-identical. 1e-12 1e-10 1e-8 1e-6 1e-4 1e-2 1 worst error over the run, as a fraction of that channel's noise = 0 1 x XNS(i) nominal nominal, vendored libm: worst 2.450e-5, inside the noise for 4.000 h nominal, platform libm: worst exactly 0, inside the noise for 4.000 h IDV(1) IDV(1), vendored libm: worst 1.600e-9, inside the noise for 4.000 h IDV(1), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(2) IDV(2), vendored libm: worst 2.660e-9, inside the noise for 4.000 h IDV(2), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(3) IDV(3), vendored libm: worst 2.490e-9, inside the noise for 4.000 h IDV(3), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(4) IDV(4), vendored libm: worst 2.120e-11, inside the noise for 4.000 h IDV(4), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(5) IDV(5), vendored libm: worst 4.160e-10, inside the noise for 4.000 h IDV(5), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(6) IDV(6), vendored libm: worst 1.490e-10, inside the noise for 4.000 h IDV(6), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(7) IDV(7), vendored libm: worst 5.120e-11, inside the noise for 4.000 h IDV(7), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(8) IDV(8), vendored libm: worst 1.420e-8, inside the noise for 4.000 h IDV(8), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(9) IDV(9), vendored libm: worst 1.310e-9, inside the noise for 4.000 h IDV(9), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(10) IDV(10), vendored libm: worst 3.660e-8, inside the noise for 4.000 h IDV(10), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(11) IDV(11), vendored libm: worst 2.410e-9, inside the noise for 4.000 h IDV(11), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(12) IDV(12), vendored libm: worst 4.870e-10, inside the noise for 4.000 h IDV(12), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(13) IDV(13), vendored libm: worst 2.070e-8, inside the noise for 4.000 h IDV(13), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(14) IDV(14), vendored libm: worst 2.450e-5, inside the noise for 4.000 h IDV(14), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(15) IDV(15), vendored libm: worst 2.450e-5, inside the noise for 4.000 h IDV(15), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(16) IDV(16), vendored libm: worst 1.170e-9, inside the noise for 4.000 h IDV(16), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(17) IDV(17), vendored libm: worst 1.210e-9, inside the noise for 4.000 h IDV(17), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(18) IDV(18), vendored libm: worst 2.420e-9, inside the noise for 4.000 h IDV(18), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(19) IDV(19), vendored libm: worst 2.450e-5, inside the noise for 4.000 h IDV(19), platform libm: worst exactly 0, inside the noise for 4.000 h IDV(20) IDV(20), vendored libm: worst 5.780e-8, inside the noise for 4.000 h IDV(20), platform libm: worst exactly 0, inside the noise for 4.000 h vendored libm vendored libm platform libm: the same exp gfortran calls worst anywhere: 2.450e-5 The shaded band is where the disagreement would be as large as the channel's own measurement noise.
All twenty-one scenarios, in units of instrument noise. The worst disagreement anywhere in each run, divided by the noise standard deviation XNS(i) of the channel it happened on, so the reference line is not zero but the point at which the plant's own instruments could resolve the difference. Filled dots are the vendored libm; rings are the platform one, and all twenty-one of those sit in the exactly-zero lane. The claim is false if a marker reaches the shaded band at one, and the explanation is false if a ring ever leaves the zero lane: that would mean the divergence is something other than transcendental rounding. Written by cargo xtask validate --tiers 4 from that run's own output, at the sweep's own horizon of four hours rather than the eight-hour nominal run tabulated above; the same sweep is recorded for B-0034 in LOG.org, and the generated index names the run that drew this copy.

The explanation is better than the one asked for. With identical transcendentals the two trajectories are not merely close, they are bit-identical for 28,800 steps on all 41 measurements. So the divergence is transcendental rounding and nothing else, and the residual under the vendored libm is twenty thousand times below what the instruments could resolve after eight hours.

Closed loop, 48 hours, 172,800 steps, nominal scenario with the driver's forced IDV(12) from hour eight, from B-0041:

libmworst XMEASworst XMV (fraction of range)within XNSat the end
platform0048.000 h0
vendored1.705e-10 at XMEAS(1)1.246e-11 at XMV(10)48.000 h7.882e-11 x XNS(13)

Bit-identical for all 172,800 closed-loop steps under the platform libm. Not one measurement, not one valve, not one step. Under the vendored libm the error never reaches a tenth of a billionth of any instrument's noise over two simulated days.

Compare the open-loop figure above: nominal at 8 hours ends at 5.149e-5 of XNS, six orders of magnitude worse. Closing the loop suppresses the amplification, which is what a controller pulling toward a setpoint should do and is worth having measured rather than assumed.

The same run measures what the control layer is actually buying. Open loop from the same start, with the valves held, trips at step 11,017 (3.060 h) on reactor pressure high. So "48 hours without tripping" is a statement about the control layer, not about a placid plant. B-0052 reproduced that trip from the public API and the CLI, at exactly the same step, without either being told about the other.

Tier 5: statistical equivalence, the real gate

Tier 5 tests equivalence rather than difference, because a failure to reject the null of no difference is not evidence of equivalence and two one-sided tests are. Every statistic ships in Rust, in tepsim-stats, with known-answer tests: Welch's t and the TOST wrapper, the two-sample Kolmogorov-Smirnov statistic, energy distance, autocorrelation, Welch power spectra and the Pearson correlation matrix. Nothing calls out to numpy or scipy.

The harness

From B-0047a: a 48-hour run produces 960 samples of 53 variables, a port run takes 766 ms in release, and a full battery of 2100 runs takes about 27 minutes per source. Over three scenarios the two sources agree bit-identically (0) under the platform libm and to 2.056e-13 under the vendored one, while five nominal seeds spread by 1.366e1, which is the scale that makes those first two numbers meaningful.

The same entry checked that the scenarios are not vacuous. All twenty disturbances move the plant within two hours; the smallest departures are IDV(3) at 4.076e-3, IDV(9) at 3.531e-3 and IDV(15) at 9.839e-3, and the largest are IDV(6) at 3.467e0 and IDV(1) at 2.456e0.

The battery

Four of the six statistics have margins that are measured rather than chosen (B-0047b). The reference's seeds are split in half, the statistic is computed against itself over twenty deterministic splits, and the cross-source value is treated as one more draw from that null, giving p = (1 + #{within >= cross}) / (K + 1) gated at 0.05. With twenty splits the cross-source value fails exactly when it is the strict maximum of the twenty-one.

The smoke battery in the CI gate is 3 scenarios, 4 seeds, 2 hours, 12 runs per source, 14 seconds. The full battery is partial: 3 of 21 scenarios at 100 seeds by 48 hours, stopped by direction after about 25 minutes.

scenarioworst mean powerFrobenius, crosswithin, maxp
nominal1.591.186e-27.5951.0000
IDV(1)1.617.080e-35.7031.0000
IDV(2)0.232.827e-51.9241.0000

Every calibrated statistic passed at p = 1.0000 on all three: the cross-source value was never the maximum of its null. The correlation matrix's Frobenius distance is three to five orders of magnitude inside the within-source spread, which matters because that matrix is exactly what a PCA-based detector consumes.

That relationship, rather than any single number, is what Tier 5 claims, so it is worth drawing:

GENERATED by `cargo xtask validate --tiers 5 --smoke` from commit `03ec3d6-dirty`. Do not edit by hand: the next run overwrites it. Tier 5: the cross-source difference against the reference's own spreadA logarithmic scatter plot. The horizontal axis is a statistic measured between the Fortran and the Rust port; the vertical axis is the same statistic measured between two halves of the Fortran's own runs. Both axes share one scale, so the diagonal is the line of equality. 3 of 5 points lie above it, meaning the two implementations differ by less than the reference differs from itself. Tier 5: the two sources against the reference's own run-to-run spread 3 scenarios x 4 seeds x 2 h, 40 samples per run, vendored libm. A point above the diagonal differs across sources by less than the reference differs from itself. 1e-12 1e-10 1e-8 1e-6 1e-4 1e-2 1 1e2 1e4 1e-12 1e-10 1e-8 1e-6 1e-4 1e-2 1 1e2 1e4 equal: the port is as different as a rerun measured between the Fortran and the port measured inside the Fortran alone KS, one variable shifted 10 sd: across sources 1.000e0, inside the Fortran 7.500e-2 KS, one variable shifted 10 sd energy, one variable shifted 10 sd: across sources 1.228e2, inside the Fortran 3.036e-2 energy, one variable shifted 10 sd correlation matrix, nominal: across sources 6.973e-12, inside the Fortran 1.091e1 correlation matrix, nominal correlation matrix, IDV(1): across sources 3.586e-12, inside the Fortran 9.873e0 correlation matrix, IDV(1) correlation matrix, IDV(2): across sources 4.248e-12, inside the Fortran 1.279e1 correlation matrix, IDV(2) inside the null above the diagonal: inside the reference's own spread outside the null below it: a difference the calibration can see The points below the diagonal are the battery's own positive control: one variable of the reference shifted by ten standard deviations and compared with the unshifted reference.
The two sources against the reference's own run-to-run spread. Horizontally, a statistic measured between the Fortran and the port. Vertically, the same statistic measured between two halves of the Fortran's own runs, which is the null the battery calibrates against. Both axes share one scale, so the diagonal is exactly the line of equality: a point above it means the two implementations differ from each other by less than the reference differs from itself. The claim is false if a cross-source point crosses to the low side. The two points that already sit there are the battery's own positive control, where one variable of the reference was shifted by ten standard deviations before comparing, and they show what a real difference looks like on the same axes. Written by cargo xtask validate --tiers 5 --smoke, so the points are the smoke battery's, 3 scenarios by 4 seeds by 2 h and fourteen seconds of running, not the full one: the figure's own subtitle says which, and the full-battery numbers are the ones tabulated above and recorded for B-0047b in LOG.org. At smoke size the permutation tests cannot reject at all, which is why the picture shows the gap between the cross-source value and its null rather than a verdict.

That plot answers "are the two sources closer to each other than the reference is to itself", which is the calibrated question. It does not answer "how much room is left", and neither does a TOST verdict: a p-value reports that a test did not reject, not by what distance. The margin is a stated quantity, a tenth of the reference's standard deviation for each variable, so the distance can simply be drawn.

GENERATED by `cargo xtask validate --tiers 5 --smoke` from commit `03ec3d6-dirty`. Do not edit by hand: the next run overwrites it. Tier 5: each variable's mean difference against its equivalence marginA logarithmic strip plot with one row per scenario. Each marker is one of the 53 variables: the paired difference between the Fortran's mean and the port's, divided by that variable's own equivalence margin. One is the gate, and the region past it is shaded. 0 of 156 gated markers reach it. 3 are exactly zero. A further 3 variables are not judged this way and are drawn hollow; a constant reference has no margin, so it sits in the zero lane. Tier 5: how far inside its margin each variable sits 3 scenarios x 4 seeds x 2 h, 40 samples per run, vendored libm, 53 variables per scenario. Paired mean difference divided by that variable's equivalence margin. 1e-30 1e-27 1e-24 1e-21 1e-18 1e-15 1e-12 1e-9 1e-6 1e-3 1 paired mean difference, as a fraction of the equivalence margin = 0 the margin nominal nominal, XMEAS/XMV column 1: difference -3.969e-15, margin 2.497e-3, ratio 1.589e-12 (judged against the margin) nominal, XMEAS/XMV column 2: difference 1.364e-12, margin 3.554e0, ratio 3.839e-13 (judged against the margin) nominal, XMEAS/XMV column 3: difference 2.501e-12, margin 3.354e0, ratio 7.458e-13 (judged against the margin) nominal, XMEAS/XMV column 4: difference 7.994e-15, margin 7.635e-3, ratio 1.047e-12 (judged against the margin) nominal, XMEAS/XMV column 5: difference 6.217e-15, margin 2.142e-2, ratio 2.902e-13 (judged against the margin) nominal, XMEAS/XMV column 6: difference -1.776e-14, margin 2.214e-2, ratio 8.022e-13 (judged against the margin) nominal, XMEAS/XMV column 7: difference 7.958e-13, margin 3.860e-1, ratio 2.062e-12 (judged against the margin) nominal, XMEAS/XMV column 8: difference 2.132e-14, margin 4.297e-2, ratio 4.961e-13 (judged against the margin) nominal, XMEAS/XMV column 9: difference -3.553e-15, margin 1.858e-3, ratio 1.912e-12 (judged against the margin) nominal, XMEAS/XMV column 10: difference -1.263e-15, margin 1.131e-3, ratio 1.116e-12 (judged against the margin) nominal, XMEAS/XMV column 11: difference -3.553e-15, margin 1.515e-2, ratio 2.345e-13 (judged against the margin) nominal, XMEAS/XMV column 12: difference 7.105e-15, margin 1.039e-1, ratio 6.837e-14 (judged against the margin) nominal, XMEAS/XMV column 13: difference 6.821e-13, margin 4.006e-1, ratio 1.703e-12 (judged against the margin) nominal, XMEAS/XMV column 14: difference -9.770e-15, margin 1.069e-1, ratio 9.137e-14 (judged against the margin) nominal, XMEAS/XMV column 15: difference 1.776e-14, margin 9.759e-2, ratio 1.820e-13 (judged against the margin) nominal, XMEAS/XMV column 16: difference 3.411e-13, margin 3.617e-1, ratio 9.429e-13 (judged against the margin) nominal, XMEAS/XMV column 17: difference 8.882e-16, margin 6.300e-2, ratio 1.410e-14 (judged against the margin) nominal, XMEAS/XMV column 18: difference -3.197e-14, margin 9.227e-3, ratio 3.465e-12 (judged against the margin) nominal, XMEAS/XMV column 19: difference -8.527e-14, margin 3.553e-1, ratio 2.400e-13 (judged against the margin) nominal, XMEAS/XMV column 20: difference 2.416e-13, margin 5.518e-2, ratio 4.379e-12 (judged against the margin) nominal, XMEAS/XMV column 21: difference -1.421e-14, margin 1.125e-2, ratio 1.264e-12 (judged against the margin) nominal, XMEAS/XMV column 22: difference -1.066e-14, margin 2.295e-2, ratio 4.644e-13 (judged against the margin) nominal, XMEAS/XMV column 23: difference 1.954e-14, margin 2.542e-2, ratio 7.686e-13 (judged against the margin) nominal, XMEAS/XMV column 24: difference -5.329e-15, margin 9.114e-3, ratio 5.847e-13 (judged against the margin) nominal, XMEAS/XMV column 25: difference 1.776e-15, margin 2.438e-2, ratio 7.285e-14 (judged against the margin) nominal, XMEAS/XMV column 26: difference -1.332e-15, margin 1.121e-2, ratio 1.188e-13 (judged against the margin) nominal, XMEAS/XMV column 27: difference -2.043e-14, margin 2.517e-2, ratio 8.115e-13 (judged against the margin) nominal, XMEAS/XMV column 28: difference -2.220e-15, margin 2.574e-3, ratio 8.625e-13 (judged against the margin) nominal, XMEAS/XMV column 29: difference 3.020e-14, margin 2.657e-2, ratio 1.137e-12 (judged against the margin) nominal, XMEAS/XMV column 30: difference -1.066e-14, margin 1.011e-2, ratio 1.055e-12 (judged against the margin) nominal, XMEAS/XMV column 31: difference 1.421e-14, margin 2.702e-2, ratio 5.260e-13 (judged against the margin) nominal, XMEAS/XMV column 32: difference -3.331e-16, margin 9.541e-3, ratio 3.491e-14 (judged against the margin) nominal, XMEAS/XMV column 33: difference -2.931e-14, margin 2.357e-2, ratio 1.244e-12 (judged against the margin) nominal, XMEAS/XMV column 34: difference -2.220e-15, margin 2.600e-3, ratio 8.541e-13 (judged against the margin) nominal, XMEAS/XMV column 35: difference -2.887e-15, margin 5.031e-3, ratio 5.738e-13 (judged against the margin) nominal, XMEAS/XMV column 36: difference -1.332e-15, margin 4.308e-3, ratio 3.093e-13 (judged against the margin) nominal, XMEAS/XMV column 37: difference -3.730e-17, margin 9.354e-4, ratio 3.987e-14 (judged against the margin) nominal, XMEAS/XMV column 38: difference -1.138e-15, margin 9.852e-4, ratio 1.155e-12 (judged against the margin) nominal, XMEAS/XMV column 39: difference -1.076e-16, margin 8.378e-4, ratio 1.284e-13 (judged against the margin) nominal, XMEAS/XMV column 40: difference 5.329e-15, margin 4.061e-2, ratio 1.312e-13 (judged against the margin) nominal, XMEAS/XMV column 41: difference 3.553e-15, margin 4.198e-2, ratio 8.463e-14 (judged against the margin) nominal, XMEAS/XMV column 42: difference 1.243e-14, margin 6.154e-2, ratio 2.020e-13 (judged against the margin) nominal, XMEAS/XMV column 43: difference 2.665e-14, margin 4.072e-2, ratio 6.544e-13 (judged against the margin) nominal, XMEAS/XMV column 44: difference -3.872e-13, margin 2.464e-1, ratio 1.572e-12 (judged against the margin) nominal, XMEAS/XMV column 45: difference 3.908e-14, margin 1.162e-1, ratio 3.364e-13 (judged against the margin) nominal, XMEAS/XMV column 46: difference 8.438e-14, margin 2.702e-2, ratio 3.123e-12 (judged against the margin) nominal, XMEAS/XMV column 47: difference -1.705e-13, margin 1.359e-1, ratio 1.255e-12 (judged against the margin) nominal, XMEAS/XMV column 48: difference -7.105e-15, margin 2.878e-1, ratio 2.469e-14 (judged against the margin) nominal, XMEAS/XMV column 49: difference -1.776e-15, margin 2.174e-1, ratio 8.172e-15 (judged against the margin) nominal, XMEAS/XMV column 50: difference -3.375e-14, margin 8.273e-2, ratio 4.079e-13 (judged against the margin) nominal, XMEAS/XMV column 51: difference 8.882e-15, margin 4.854e-2, ratio 1.830e-13 (judged against the margin) nominal, XMEAS/XMV column 52: difference 1.421e-14, margin 1.399e-1, ratio 1.016e-13 (judged against the margin) nominal, XMEAS/XMV column 53: difference exactly 0, margin exactly 0, no ratio: the margin is zero (reference never moved: no mean to compare) IDV(1) IDV(1), XMEAS/XMV column 1: difference -5.551e-16, margin 2.005e-2, ratio 2.768e-14 (judged against the margin) IDV(1), XMEAS/XMV column 2: difference 1.933e-12, margin 3.633e0, ratio 5.319e-13 (judged against the margin) IDV(1), XMEAS/XMV column 3: difference 3.183e-12, margin 3.713e0, ratio 8.574e-13 (judged against the margin) IDV(1), XMEAS/XMV column 4: difference 2.043e-14, margin 3.042e-2, ratio 6.715e-13 (judged against the margin) IDV(1), XMEAS/XMV column 5: difference 1.421e-14, margin 2.145e-2, ratio 6.625e-13 (judged against the margin) IDV(1), XMEAS/XMV column 6: difference 3.375e-14, margin 2.306e-2, ratio 1.463e-12 (judged against the margin) IDV(1), XMEAS/XMV column 7: difference 1.705e-12, margin 3.978e0, ratio 4.287e-13 (judged against the margin) IDV(1), XMEAS/XMV column 8: difference -7.461e-14, margin 1.051e-1, ratio 7.100e-13 (judged against the margin) IDV(1), XMEAS/XMV column 9: difference -3.553e-15, margin 1.973e-3, ratio 1.801e-12 (judged against the margin) IDV(1), XMEAS/XMV column 10: difference 2.498e-16, margin 3.019e-3, ratio 8.275e-14 (judged against the margin) IDV(1), XMEAS/XMV column 11: difference -4.619e-14, margin 1.125e-1, ratio 4.107e-13 (judged against the margin) IDV(1), XMEAS/XMV column 12: difference 2.309e-14, margin 1.041e-1, ratio 2.218e-13 (judged against the margin) IDV(1), XMEAS/XMV column 13: difference 1.364e-12, margin 3.968e0, ratio 3.438e-13 (judged against the margin) IDV(1), XMEAS/XMV column 14: difference -6.217e-15, margin 1.068e-1, ratio 5.823e-14 (judged against the margin) IDV(1), XMEAS/XMV column 15: difference 1.243e-14, margin 9.910e-2, ratio 1.255e-13 (judged against the margin) IDV(1), XMEAS/XMV column 16: difference 2.160e-12, margin 4.180e0, ratio 5.168e-13 (judged against the margin) IDV(1), XMEAS/XMV column 17: difference exactly 0, margin 6.513e-2, ratio exactly 0 (judged against the margin) IDV(1), XMEAS/XMV column 18: difference -3.197e-14, margin 4.795e-2, ratio 6.668e-13 (judged against the margin) IDV(1), XMEAS/XMV column 19: difference 1.350e-13, margin 9.887e-1, ratio 1.365e-13 (judged against the margin) IDV(1), XMEAS/XMV column 20: difference 1.705e-13, margin 2.542e-1, ratio 6.708e-13 (judged against the margin) IDV(1), XMEAS/XMV column 21: difference -6.040e-14, margin 1.362e-2, ratio 4.433e-12 (judged against the margin) IDV(1), XMEAS/XMV column 22: difference exactly 0, margin 7.191e-2, ratio exactly 0 (judged against the margin) IDV(1), XMEAS/XMV column 23: difference 6.217e-15, margin 8.721e-2, ratio 7.129e-14 (judged against the margin) IDV(1), XMEAS/XMV column 24: difference -3.109e-15, margin 1.481e-2, ratio 2.099e-13 (judged against the margin) IDV(1), XMEAS/XMV column 25: difference 3.819e-14, margin 1.329e-1, ratio 2.873e-13 (judged against the margin) IDV(1), XMEAS/XMV column 26: difference -3.775e-15, margin 1.143e-2, ratio 3.302e-13 (judged against the margin) IDV(1), XMEAS/XMV column 27: difference -3.464e-14, margin 2.711e-2, ratio 1.278e-12 (judged against the margin) IDV(1), XMEAS/XMV column 28: difference -1.554e-15, margin 4.351e-3, ratio 3.572e-13 (judged against the margin) IDV(1), XMEAS/XMV column 29: difference 6.217e-15, margin 1.415e-1, ratio 4.393e-14 (judged against the margin) IDV(1), XMEAS/XMV column 30: difference -8.882e-16, margin 1.873e-2, ratio 4.743e-14 (judged against the margin) IDV(1), XMEAS/XMV column 31: difference 5.151e-14, margin 2.158e-1, ratio 2.387e-13 (judged against the margin) IDV(1), XMEAS/XMV column 32: difference -2.665e-15, margin 9.819e-3, ratio 2.714e-13 (judged against the margin) IDV(1), XMEAS/XMV column 33: difference -4.086e-14, margin 3.639e-2, ratio 1.123e-12 (judged against the margin) IDV(1), XMEAS/XMV column 34: difference -2.554e-15, margin 5.717e-3, ratio 4.467e-13 (judged against the margin) IDV(1), XMEAS/XMV column 35: difference -8.216e-15, margin 2.191e-2, ratio 3.749e-13 (judged against the margin) IDV(1), XMEAS/XMV column 36: difference -3.331e-15, margin 1.209e-2, ratio 2.756e-13 (judged against the margin) IDV(1), XMEAS/XMV column 37: difference -4.597e-17, margin 9.403e-4, ratio 4.889e-14 (judged against the margin) IDV(1), XMEAS/XMV column 38: difference -1.249e-15, margin 3.665e-3, ratio 3.407e-13 (judged against the margin) IDV(1), XMEAS/XMV column 39: difference -1.249e-16, margin 8.646e-4, ratio 1.445e-13 (judged against the margin) IDV(1), XMEAS/XMV column 40: difference -3.553e-15, margin 4.165e-2, ratio 8.529e-14 (judged against the margin) IDV(1), XMEAS/XMV column 41: difference 3.553e-15, margin 4.252e-2, ratio 8.356e-14 (judged against the margin) IDV(1), XMEAS/XMV column 42: difference 3.375e-14, margin 6.168e-2, ratio 5.472e-13 (judged against the margin) IDV(1), XMEAS/XMV column 43: difference 4.263e-14, margin 4.234e-2, ratio 1.007e-12 (judged against the margin) IDV(1), XMEAS/XMV column 44: difference -6.217e-14, margin 1.971e0, ratio 3.154e-14 (judged against the margin) IDV(1), XMEAS/XMV column 45: difference 1.243e-13, margin 2.137e-1, ratio 5.819e-13 (judged against the margin) IDV(1), XMEAS/XMV column 46: difference 4.086e-14, margin 8.073e-2, ratio 5.061e-13 (judged against the margin) IDV(1), XMEAS/XMV column 47: difference -8.882e-15, margin 3.736e-1, ratio 2.377e-14 (judged against the margin) IDV(1), XMEAS/XMV column 48: difference -1.421e-14, margin 2.890e-1, ratio 4.917e-14 (judged against the margin) IDV(1), XMEAS/XMV column 49: difference -1.421e-14, margin 2.198e-1, ratio 6.464e-14 (judged against the margin) IDV(1), XMEAS/XMV column 50: difference -1.954e-14, margin 1.935e-1, ratio 1.010e-13 (judged against the margin) IDV(1), XMEAS/XMV column 51: difference 1.954e-13, margin 5.172e-2, ratio 3.778e-12 (judged against the margin) IDV(1), XMEAS/XMV column 52: difference 1.066e-14, margin 1.539e-1, ratio 6.924e-14 (judged against the margin) IDV(1), XMEAS/XMV column 53: difference exactly 0, margin exactly 0, no ratio: the margin is zero (reference never moved: no mean to compare) IDV(2) IDV(2), XMEAS/XMV column 1: difference -4.302e-16, margin 2.562e-3, ratio 1.679e-13 (judged against the margin) IDV(2), XMEAS/XMV column 2: difference -1.137e-12, margin 3.680e0, ratio 3.089e-13 (judged against the margin) IDV(2), XMEAS/XMV column 3: difference -2.274e-12, margin 3.573e0, ratio 6.364e-13 (judged against the margin) IDV(2), XMEAS/XMV column 4: difference 3.553e-15, margin 9.253e-3, ratio 3.839e-13 (judged against the margin) IDV(2), XMEAS/XMV column 5: difference 1.954e-14, margin 2.143e-2, ratio 9.118e-13 (judged against the margin) IDV(2), XMEAS/XMV column 6: difference 1.066e-14, margin 2.405e-2, ratio 4.432e-13 (judged against the margin) IDV(2), XMEAS/XMV column 7: difference -2.274e-13, margin 1.015e0, ratio 2.239e-13 (judged against the margin) IDV(2), XMEAS/XMV column 8: difference -5.684e-14, margin 4.790e-2, ratio 1.187e-12 (judged against the margin) IDV(2), XMEAS/XMV column 9: difference -1.066e-14, margin 1.907e-3, ratio 5.590e-12 (judged against the margin) IDV(2), XMEAS/XMV column 10: difference -6.106e-16, margin 6.669e-3, ratio 9.156e-14 (judged against the margin) IDV(2), XMEAS/XMV column 11: difference 7.105e-15, margin 1.801e-2, ratio 3.945e-13 (judged against the margin) IDV(2), XMEAS/XMV column 12: difference -3.553e-15, margin 1.039e-1, ratio 3.420e-14 (judged against the margin) IDV(2), XMEAS/XMV column 13: difference 1.137e-13, margin 1.025e0, ratio 1.109e-13 (judged against the margin) IDV(2), XMEAS/XMV column 14: difference 1.776e-15, margin 1.069e-1, ratio 1.661e-14 (judged against the margin) IDV(2), XMEAS/XMV column 15: difference 2.842e-14, margin 9.763e-2, ratio 2.911e-13 (judged against the margin) IDV(2), XMEAS/XMV column 16: difference -6.821e-13, margin 1.009e0, ratio 6.758e-13 (judged against the margin) IDV(2), XMEAS/XMV column 17: difference 3.553e-15, margin 6.300e-2, ratio 5.639e-14 (judged against the margin) IDV(2), XMEAS/XMV column 18: difference 1.421e-14, margin 1.158e-2, ratio 1.227e-12 (judged against the margin) IDV(2), XMEAS/XMV column 19: difference -1.137e-13, margin 3.842e-1, ratio 2.959e-13 (judged against the margin) IDV(2), XMEAS/XMV column 20: difference 7.105e-14, margin 1.103e-1, ratio 6.444e-13 (judged against the margin) IDV(2), XMEAS/XMV column 21: difference 7.105e-15, margin 1.261e-2, ratio 5.633e-13 (judged against the margin) IDV(2), XMEAS/XMV column 22: difference 2.487e-14, margin 2.757e-2, ratio 9.021e-13 (judged against the margin) IDV(2), XMEAS/XMV column 23: difference -1.776e-15, margin 2.537e-2, ratio 7.002e-14 (judged against the margin) IDV(2), XMEAS/XMV column 24: difference -1.332e-15, margin 2.779e-2, ratio 4.794e-14 (judged against the margin) IDV(2), XMEAS/XMV column 25: difference 1.776e-15, margin 2.538e-2, ratio 6.999e-14 (judged against the margin) IDV(2), XMEAS/XMV column 26: difference -2.220e-16, margin 1.144e-2, ratio 1.940e-14 (judged against the margin) IDV(2), XMEAS/XMV column 27: difference -3.553e-15, margin 2.675e-2, ratio 1.328e-13 (judged against the margin) IDV(2), XMEAS/XMV column 28: difference -1.110e-16, margin 3.105e-3, ratio 3.575e-14 (judged against the margin) IDV(2), XMEAS/XMV column 29: difference 5.329e-15, margin 2.755e-2, ratio 1.934e-13 (judged against the margin) IDV(2), XMEAS/XMV column 30: difference -3.109e-15, margin 4.138e-2, ratio 7.512e-14 (judged against the margin) IDV(2), XMEAS/XMV column 31: difference -8.882e-16, margin 2.903e-2, ratio 3.059e-14 (judged against the margin) IDV(2), XMEAS/XMV column 32: difference -1.110e-16, margin 9.546e-3, ratio 1.163e-14 (judged against the margin) IDV(2), XMEAS/XMV column 33: difference 3.944e-31, margin 2.819e-2, ratio 1.399e-29 (judged against the margin) IDV(2), XMEAS/XMV column 34: difference 6.661e-16, margin 3.329e-3, ratio 2.001e-13 (judged against the margin) IDV(2), XMEAS/XMV column 35: difference 1.554e-15, margin 5.579e-3, ratio 2.786e-13 (judged against the margin) IDV(2), XMEAS/XMV column 36: difference 1.665e-15, margin 4.433e-3, ratio 3.756e-13 (judged against the margin) IDV(2), XMEAS/XMV column 37: difference -5.204e-18, margin 9.349e-4, ratio 5.566e-15 (judged against the margin) IDV(2), XMEAS/XMV column 38: difference 1.665e-16, margin 1.004e-3, ratio 1.659e-13 (judged against the margin) IDV(2), XMEAS/XMV column 39: difference 1.388e-17, margin 8.411e-4, ratio 1.650e-14 (judged against the margin) IDV(2), XMEAS/XMV column 40: difference exactly 0, margin 4.051e-2, ratio exactly 0 (judged against the margin) IDV(2), XMEAS/XMV column 41: difference 5.329e-15, margin 4.173e-2, ratio 1.277e-13 (judged against the margin) IDV(2), XMEAS/XMV column 42: difference -1.243e-14, margin 6.259e-2, ratio 1.987e-13 (judged against the margin) IDV(2), XMEAS/XMV column 43: difference -3.020e-14, margin 4.172e-2, ratio 7.239e-13 (judged against the margin) IDV(2), XMEAS/XMV column 44: difference -3.642e-14, margin 2.522e-1, ratio 1.444e-13 (judged against the margin) IDV(2), XMEAS/XMV column 45: difference 1.066e-14, margin 1.262e-1, ratio 8.443e-14 (judged against the margin) IDV(2), XMEAS/XMV column 46: difference -1.243e-14, margin 3.655e-2, ratio 3.402e-13 (judged against the margin) IDV(2), XMEAS/XMV column 47: difference -7.105e-14, margin 7.709e-1, ratio 9.218e-14 (judged against the margin) IDV(2), XMEAS/XMV column 48: difference 5.329e-15, margin 2.874e-1, ratio 1.854e-14 (judged against the margin) IDV(2), XMEAS/XMV column 49: difference 7.105e-15, margin 2.173e-1, ratio 3.270e-14 (judged against the margin) IDV(2), XMEAS/XMV column 50: difference -2.665e-14, margin 8.568e-2, ratio 3.110e-13 (judged against the margin) IDV(2), XMEAS/XMV column 51: difference 7.994e-14, margin 5.097e-2, ratio 1.568e-12 (judged against the margin) IDV(2), XMEAS/XMV column 52: difference -7.889e-31, margin 1.399e-1, ratio 5.637e-30 (judged against the margin) IDV(2), XMEAS/XMV column 53: difference exactly 0, margin exactly 0, no ratio: the margin is zero (reference never moved: no mean to compare) judged against the margin judged against the margin not judged this way: a constant reference, or a stuck valve worst gated: 5.590e-12 A marker past the shaded line would be a variable whose mean moved by more than the margin allows.
Each variable's mean difference against its own equivalence margin. One marker per variable per scenario: the paired difference between the Fortran's mean and the port's, divided by that variable's equivalence margin. One is the gate and the region past it is shaded, so the claim is false the moment a solid marker reaches the shaded band. Nothing is close: the worst gated variable sits at 5.590e-12 of its margin, which is about eleven orders of magnitude of headroom. Hollow markers are the variables the moment gate does not apply to, drawn rather than dropped, because omitting them would quietly shrink the denominator a reader counts against. A constant reference has no margin at all and so has no ratio, and sits in the zero lane; a valve stuck by the scenario's own fault is judged on its distribution instead, which is the decision recorded as B-0047d. Written by cargo xtask validate --tiers 5 --smoke, at the same smoke size as the figure above and for the same reason.

TOST power against battery size, measured in the same entry:

batteryworst power
4 seeds, 2 h (smoke)11.93
8 seeds, 12 h3.38
8 seeds, 48 h1.04
100 seeds, 48 h, nominal1.59
100 seeds, 48 h, IDV(2)0.23

A power above 1 means the battery is underpowered for the margin, not that the sources differ. PLAN.org sets the mean margin at a tenth of the pooled standard deviation, and a disturbed plant has a much larger pooled spread than a quiet one, so the same absolute run-to-run wander sits comfortably inside the margin under IDV(2) and outside it at the nominal operating point. At the nominal plant the margin needs about 255 seeds, not 100. The variables that run out of power are the manipulated ones: an integrating controller's output has a random-walk component, so its mean over 48 hours varies between seeds by much more than a tenth of its own within-run spread. That is a fact about the plant, not about the port, and the battery reports it as undecided rather than as a difference while still asserting on the measured gap.

Physics invariants

These are the only tests in the whole ladder that can catch an error the port faithfully inherited, because every other tier compares against the Fortran and an error in the Fortran is invisible to all of them. From B-0046:

invariantFortranportgate
I-1 reaction mass, 200 states4.664e-163.498e-161e-14
I-2 plant mass balance, nominal2.079e-162.344e-161e-13
I-2 along 2 h open loop6.556e-16not run1e-13
I-3 inert reaction termexactly 0exactly 0exact
I-4 per-component moles, 1 h7.436e-152.892e-151e-13

All four reactions balance exactly with the published molecular weights, `2 + 28

  • 32 = 62, 2 + 28 + 46 = 76, 2 + 46 = 48and3 x 32 = 2 x 48`, which is why I-1 is an equality rather than a tolerance.

The invariants' teeth were checked by mutation, and one result is worth repeating because it is not obvious. Deleting the reaction term from the reactor's component balance leaves the total mass balance passing at 2e-16. That is not a weak test, it is a true fact about the invariant: I-2 is the molecular-weight-weighted sum of I-4, and I-1 says the reaction is mass-neutral, so the reaction term cancels out of I-2 exactly. Total mass conservation is blind to stoichiometry by construction. The unweighted per-component balance, I-4, sees it immediately. An invariant that is a sum of other invariants is weaker than the set, and how much weaker is not obvious from reading it.

Tiers 6 through 10

These have not run. They are listed here so that the shape of the remaining claim is visible rather than implied.

Tier 6, downstream-task equivalence, is the operational definition of "practically equivalent" for this project's research audience: train a detector suite on Fortran d00 data, evaluate on Fortran and on Rust test data, then reverse it, and compare detection rate, false alarm rate and detection delay. The claim to be able to make is that cross-source performance matches within-source performance within its own run-to-run variability. Backlog item B-0050.

Tier 7, published dataset reproduction, attempts direct reproduction of the bundled d00 through d21 files under the documented generation protocol, reporting per-file agreement with the Tier 5 machinery. The groundwork exists: teprob.f:1187-1256 carries fifty-four generator words in comments, one per published dataset, and B-0047a transcribed and asserted them. Three facts about that table are worth knowing in advance. Thirty-four of the fifty-four exceed 2^32, which is not a transcription error because TESUB7 reduces any seed on the first draw. Twenty-seven are even, and a multiplicative generator modulo a power of two keeps the factors of two its seed has, so those runs have a shorter period and low bits that never move. Reproducing the published files means doing the same rather than fixing it. Backlog item B-0051.

Tier 8 is differential fuzzing, Tier 9 is cross-platform determinism by golden BLAKE3 digest including wasm in a real browser, and Tier 10 requires that every quirk fix ship with a measured delta from the full Tier 5 battery with the fix on and off. None has started.

Two bugs that only long runs found

Worth recording because they are the argument for running the battery long rather than often (B-0047b).

TRCN, the Tier 3 trace counter, is a Fortran INTEGER that nothing clears between evaluations. At about 264 draws per step over 172,800 steps it passes 2^31 after roughly fifty runs, goes negative, passes the capacity guard, and writes outside the array. The symptom was a SIGSEGV deep inside the Fortran, tens of millions of steps into the battery, with nothing nearby to suggest why: the same seed ran perfectly in isolation, and starting at seed 40 moved the crash to seed 90, which is what identified it as cumulative rather than data-dependent. Nothing in Tiers 1 to 4 was ever affected, because no shorter run approached the overflow.

welch_t computed (n - 1) on a summary with no observations. In debug that panics; in release it wraps and returns a degrees-of-freedom figure that looks like a number. The case is XMV(12), the agitator, which no controller ever writes, so every run has zero variance and the log-variance ensemble is empty.

Both are the same lesson. A validation harness is code, it has bugs, and the bugs it has are the ones only its longest runs reach.