For a reviewer:
- Decision: gradient-boosted trees over a hand-built physics feature mart, exported to ONNX for in-browser scoring; interpretability and parity over raw model flexibility.
- Trade-off: tree models on tabular features, not deep learning, so gains come from the decomposition’s signal rather than model capacity.
- Proof: train/serve parity to
atol=1e-5, calibrated prediction intervals, and a no-leakage feature spine, all enforced in CI.
5 XGBoost models
Three question families degradation pace loss (p10/p50/p90), laps until
the cliff (4-class), and remaining stint life. One shared feature matrix,
one training engine, one pipeline.
32 features in 6 groups
Per-lap stint-position, compound-prior, cliff-prior, thermal, and
dirty-air signals all read from
fct_cliff_prediction_features,
never re-derived in ML.189 tests
28 Leakage Spine + 5 ONNX Parity + 3 Predict Schema + 19 Evaluation Gates + 5 CRPS + 3 Targets + 18 Survival + 18 Version Contract + 19 Attainable Ceilings + 33 Within-stint Attribution + 24 Fit Parity + 14 Search Space every build.
82,470 training laps
2018–2024 F1 seasons; holdout is
MAX(race_year) + 1
(no data yet the final CV fold stands in until it ingests).All five beat baseline
Every model is evaluated against a strong per-cohort non-leakage baseline
and wins. Coverage: quantile interval nominal 0.800 → empirical 0.814.
One command
make ml-all runs the full 8-stage pipeline end-to-end:
features → tune → train → evaluate → predict → onnx → card → reference.Why ML is needed on top of the physics decomposition
The statistical layer already estimates a tyre cliff for every(circuit, compound, season) cohort using Kaplan-Meier survival analysis. That is the right answer to a population question: when does the average Soft at Bahrain fall off? It cannot answer whether this stint, in these conditions, is about to drop.
The cliff is not a fixed lap number. It moves with:
- Thermal history how hard the tyre has been pushed (
push_residual, cumulative surface and bulk load) - Dirty air laps spent in another car’s wake overheat the surface (
dirty_air_thermal_load_surface/bulk) - Ambient and air density a hot, thin-air afternoon cliffs earlier than a cool evening session
- Fuel load a heavy car early in the stint loads the tyre differently from a light car at the end
- Compound generation 2018 legacy compounds behave unlike the 2019+ range
The feature source
All five models read from a single contracted mart:fct_cliff_prediction_features. The models consume 33 features across six physics-grouped families (pruned from 42 across ten by Phase 9’s noise-floor ablation, then joined by Phase 10a’s proximity group — the first features in the project sourced from the position channel rather than from lap times or the car channel) nothing is hand-engineered inside the ML layer. The physics lives in the dbt transform layer; the machine layer reads it.
One mart, two target spans. The quantile trio and the cliff classifier train on 114,270 laps that carry the degradation and cliff targets; the stint-life regressor trains on 120,934 laps the spans differ because a lap can have a usable stint-life count even where its next-lap jump is undefined. All seven seasons (2018–2024) feed in, and the 19,144-lap 2024 fold stands in as the evaluation holdout until 2025 ingests.
Three question families, five models
How much pace will the tyre lose?
Three quantile regressors p10, p50, and p90 predict
next_lap_degradation_jump_s. Together they form a calibrated 80% prediction interval: p50 is the best single guess; p10 and p90 bound it. The target is legitimately negative roughly 44% of the time a tyre coming into its working window genuinely gains pace.How close is the cliff?
A 4-class classifier predicts
laps_until_cliff_class: 0_to_2, 3_to_5, 6_plus, or none_in_stint. Most laps have no cliff ahead at all-the balance is roughly 9 / 7 / 9 / 76 percent-so the model trains with balanced class weights to keep the cliff-window classes from being drowned out by none_in_stint.How many usable laps remain?
A single regressor predicts
remaining_stint_life_laps the synthesised, non-negative count of usable laps left in the stint. This is what drives the strategy view’s “this set is done in ~N laps” readout.All five run in-browser
Every model is exported to ONNX and scored client-side via the React app no server round-trip required. ONNX parity tests confirm numbers in the browser match numbers from training to within
atol=1e-5.How the models fit into the broader system
The machine layer reads the warehouse read-only and never writes to the application layer. Reproducing the full pipeline takes a single command:Go deeper
Feature contract
The 33 features across 6 physics groups, where each comes from in the mart, and the read-only contract that keeps physics and ML cleanly separated.
The pipeline
The eight-stage
make ml-all DAG features → tune → train → evaluate → predict → onnx → card → reference and what each stage reads, writes, and guarantees.Validation & Trust
Season-grouped time-series cross-validation, the leakage spine, calibration coverage, and the
MAX+1-derived holdout policy.The CI Contract
All 189 tests 28 leakage-spine guards, 5 ONNX parity, 3 predict-schema, 19 evaluation gates, 5 CRPS, 3 target/censoring, 18 survival/AFT, 18 version-contract, 19 attainable-ceiling, 33 within-stint attribution, 24 fit-parity, 14 search-space grouped and explained.
Model Reference
The generated model card: per-model hyperparameters, headline metrics, calibration, cohorts, and reproducibility.