Clear target separation makes tyre age one of the first raw variables worth testing.
DETAILED EXPLORATORY DATA ANALYSIS
Lap-by-lap, from raw data to a baseline.
A native web translation of the project notebook—rebuilt from its recorded aggregates, model outputs, section order, and conclusions.
TARGET DISTRIBUTION
One in five laps is followed by a pit stop.
The notebook begins with the target itself. PitNextLap is imbalanced, but not severely enough to demand aggressive resampling—especially because ROC AUC measures ranking quality.
PitNextLap distribution
NUMERIC FEATURE DISTRIBUTIONS
Train and test overlap across the feature space.
The notebook compares nine shared numeric features. The native profiles below use its exact train and test means and standard deviations to show their overlap without relying on a raster image.
Train vs test statistical profiles
- Train μ
- 14.16
- Test μ
- 14.16
- Mean drift
- +0.02%
Train and test distributions heavily overlap for most features. Local validation should therefore be a useful leaderboard proxy.
FEATURE VS TARGET
The separation appears in tyre life and race timing.
The notebook's class overlays point to tyre life and race timing. Because their histogram bins were not saved in the notebook JSON, the native chart uses the exact printed target correlations as a reproducible signal summary.
Feature signal toward PitNextLap
Later race laps carry more pit-stop likelihood, although timing features overlap.
Strategy phases emerge more clearly than they do in raw LapTime_Delta.
CORRELATION HEATMAP
Strong signal, redundancy, and one inverted feature.
The relationship map preserves the notebook's most consequential feature-to-feature and feature-to-target values in a responsive, readable form.
Relationship map
LapNumber and RaceProgress correlate at 0.96; TyreLife and LapNumber at 0.65. Interaction features may be more useful than carrying every raw timing signal independently.
CATEGORICAL ANALYSIS
The compound ordering defies real-world intuition.
The notebook analyses pit rate by compound and year. These exact group means are now rendered as native bars, keeping the 2023 anomaly visible without the old Matplotlib panel.
Categorical pit rates
HARD has the highest pit rate—not SOFT. This synthetic dataset does not fully honour domain logic, so assumptions must not override the observed data.
PIT TIMING
There is no single universal pit window.
The notebook answers when stops happen by race lap, race progress, tyre age, and compound. Its raw histogram bins were not serialized, so this native view clearly labels the observed ranges as notebook-derived rather than exact bin counts.
Observed pit windows
Stops spread broadly across laps 5-55, consistent with mixed one-stop and two-stop strategies.
Activity peaks around 0.35-0.60 and tapers sharply after 0.70.
The strongest concentration is 10-20 laps; stops become rare after 30 laps.
Tyre-life distributions overlap heavily, so compound alone does not determine stint length.
DEGRADATION ANALYSIS
LapTime_Delta is inverted, not useless.
The notebook bins tyre age and prints the average LapTime_Delta for every interval. Those exact grouped means drive the native chart below.
LapTime_Delta by tyre-life bin
The raw delta is relative to a reference pace. Rolling rate-of-change should be more useful than interpreting it as direct lap-over-lap tyre degradation.
TRAIN / TEST DRIFT
Only two features move by roughly five percent.
The notebook computes mean differences feature by feature. Every numeric feature remains under ten percent drift, and most stay below half a percent.
Train vs test mean drift
No drift correction was required. Position_Change reaches 5.18% and LapTime_Delta 5.10%; all other mean differences are below 0.5%.
RACE-LEVEL PIT RATE
Circuit identity carries real strategy signal.
The notebook ranks every race by observed PitNextLap rate. Its printed top and bottom ten rows are recreated below, including Pre-Season Testing.
Pit rate by race
Recommendation from the notebook: encode Race because circuit-level pit rates vary substantially. Remove Pre-Season Testing because its stops do not follow race-strategy logic.
BASELINE LIGHTGBM
EDA decisions flow into a time-aware baseline.
The notebook closes by encoding Driver, Compound, and Race; holding out 2025; training LightGBM; evaluating AUC; inspecting feature importance; and writing a submission.
Driver · Compound · Race
346,246 rows · 2022-2024
92,894 rows · 2025
ROC AUC · year-aware
LightGBM feature importance
NOTEBOOK CONCLUSION