compound, circuit_key, and is_rain_lap dimensions. Cells where the model trails its per-cohort baseline are recorded in the model card and published here. The contract: surface the losses, do not drop them.
Underperforming cohort cells
The cell count comes from the model card it updates on everymake ml-card run. The table below reflects the most recent evaluation.
Model score and Baseline are the headline metric for each family: pinball ↓ for p10/p50/p90, macro-F1 ↑ for the classifier, RMSE ↓ for stint-life. Δ is signed so that negative always means the model trails, whichever direction its metric runs; a Δ of exactly 0 is a tie on a cell too small to separate them.
What the losses mean
The_other compound cells (n = 9, 12). The _other bucket covers legacy or unrecognised compound codes. With nine training rows, the baseline has nothing to learn it is essentially the global mean. These losses are structural artefacts of the small cell size rather than evidence the model has a problem with any real compound.
The cliff classifier’s two cells are ties, not losses. Both are degenerate: nine _other-compound rows and fifteen rain laps, where model and baseline score identically. The classifier no longer trails on INTERMEDIATE compounds or at the Miami GP, which it did before the laps_until_cliff_class label was repaired.
The is_rain_lap cells (n = 15 rain laps). Fifteen laps in the evaluation set. A low-n artefact.
Circuit-level losses (stint-life regressor). The stint-life regressor’s per-cohort baseline is near-oracle: it uses the true final stint length known only once the stint is over as its “prediction.” On circuits where stint lengths are highly predictable (the Dutch GP historically runs long stints, the Las Vegas circuit tends toward specific strategic windows), that oracle baseline is genuinely hard to beat. The model wins globally (RMSE 7.73 vs 9.40) because the baseline is not available in real-time, but in the retrospective evaluation on those circuits the oracle wins. Las Vegas is the widest cell (8.82 vs 5.30).
The losses on Canadian GP (pinball) and on Bahrain, British and Japanese GPs (stint-life) are smaller, and mix sample-size effects with track-specific behaviour the model does not fully capture.
The CI contract
test_evaluate.py asserts that cohort losses are computed and surfaced on every build, not that the model wins every cell. The gate is: “we measured this cohort, and if we lost, we know it.” A model that silently skipped the cohort evaluation would fail the gate even if it was otherwise excellent.Relationships
Validation
The global validation context where the model wins overall.
Evaluation gates
The CI tests that enforce cohort surfacing and the beats-baseline contract.
Model Reference
The full cohort table from the latest model card updated on every
make ml-card run.Tyre cliff
The physics behind the cliff concept the classifier is trying to predict.