ORIGINAL NOTEBOOK · WEB REPORT
S6E5 · Lap-by-lap analysis

DETAILED EXPLORATORY DATA ANALYSIS

Lap-by-lap, from raw data to a baseline.

A native web translation of the project notebook—rebuilt from its recorded aggregates, model outputs, section order, and conclusions.

439,140Train rows
188,165Test rows
16 / 15Train / test columns
0Missing values
53.6 MBTraining memory
01

TARGET DISTRIBUTION

One in five laps is followed by a pit stop.

The notebook begins with the target itself. PitNextLap is imbalanced, but not severely enough to demand aggressive resampling—especially because ROC AUC measures ranking quality.

Exact notebook data

PitNextLap distribution

HTML · CSS
19.9%pit next lap
Stay out351,759
Class 0 · 80.10%
Pit next lap87,381
Class 1 · 19.90%
Exact counts and proportions from notebook cell 11
02

NUMERIC FEATURE DISTRIBUTIONS

Train and test overlap across the feature space.

The notebook compares nine shared numeric features. The native profiles below use its exact train and test means and standard deviations to show their overlap without relying on a raster image.

Exact notebook data

Train vs test statistical profiles

HTML · CSS
TrainTestDot = mean · band = ±1 standard deviation
ACTIVE CHANNELTyreLife
Train μ
14.16
Test μ
14.16
Mean drift
+0.02%
Mean ± standard deviation from the notebook output; each feature uses its own scale
Main observation

Train and test distributions heavily overlap for most features. Local validation should therefore be a useful leaderboard proxy.

03

FEATURE VS TARGET

The separation appears in tyre life and race timing.

The notebook's class overlays point to tyre life and race timing. Because their histogram bins were not saved in the notebook JSON, the native chart uses the exact printed target correlations as a reproducible signal summary.

Exact notebook data

Feature signal toward PitNextLap

HTML · CSS
TyreLife
+0.274
LapNumber
+0.267
Stint
+0.198
RaceProgress
+0.185
Position_Change
+0.046
Position
+0.021
LapTime_Delta
-0.005
LapTime (s)
-0.034
Cumulative_Degradation
-0.167
Exact Pearson correlation values printed by notebook cell 20
A
TyreLife

Clear target separation makes tyre age one of the first raw variables worth testing.

B
LapNumber

Later race laps carry more pit-stop likelihood, although timing features overlap.

C
RaceProgress

Strategy phases emerge more clearly than they do in raw LapTime_Delta.

04

CORRELATION HEATMAP

Strong signal, redundancy, and one inverted feature.

The relationship map preserves the notebook's most consequential feature-to-feature and feature-to-target values in a responsive, readable form.

Exact notebook data

Relationship map

HTML · CSS
LapNumber↔RaceProgress+0.96
TyreLife↔LapNumber+0.65
TyreLife↔PitNextLap+0.274
LapNumber↔PitNextLap+0.267
Cumulative_Degradation↔PitNextLap-0.167
LapTime_Delta↔PitNextLap-0.005
Strong relationships and target correlations retained by the notebook
+0.27TyreLife
+0.27LapNumber
+0.20Stint
+0.19RaceProgress
-0.17Cumulative_Degradation
-0.005LapTime_Delta
Redundancy warning

LapNumber and RaceProgress correlate at 0.96; TyreLife and LapNumber at 0.65. Interaction features may be more useful than carrying every raw timing signal independently.

05

CATEGORICAL ANALYSIS

The compound ordering defies real-world intuition.

The notebook analyses pit rate by compound and year. These exact group means are now rendered as native bars, keeping the 2023 anomaly visible without the old Matplotlib panel.

Exact notebook data

Categorical pit rates

HTML · CSS

Compound

HARD32.8%
SOFT19.3%
INTERMEDIATE15.2%
MEDIUM10.1%
WET2.5%

Season

202429.5%
202528.4%
202226.7%
20231.0%
Exact group means from notebook cell 23
HARD 32.8%SOFT 19.3%INTERMEDIATE 15.2%MEDIUM 10.1%WET 2.5%
Notebook warning

HARD has the highest pit rate—not SOFT. This synthetic dataset does not fully honour domain logic, so assumptions must not override the observed data.

06

PIT TIMING

There is no single universal pit window.

The notebook answers when stops happen by race lap, race progress, tyre age, and compound. Its raw histogram bins were not serialized, so this native view clearly labels the observed ranges as notebook-derived rather than exact bin counts.

Notebook-derived

Observed pit windows

HTML · CSS
Race lap5–55
0broad activity78
Race progress35–60
0strongest phase100
Tyre age10–20
0highest concentration50
After 30 tyre laps30–50
0events become rare50
Ranges stated in the notebook analysis; histogram bin arrays were not serialized
01
LapNumber

Stops spread broadly across laps 5-55, consistent with mixed one-stop and two-stop strategies.

02
RaceProgress

Activity peaks around 0.35-0.60 and tapers sharply after 0.70.

03
TyreLife

The strongest concentration is 10-20 laps; stops become rare after 30 laps.

04
By compound

Tyre-life distributions overlap heavily, so compound alone does not determine stint length.

07

DEGRADATION ANALYSIS

LapTime_Delta is inverted, not useless.

The notebook bins tyre age and prints the average LapTime_Delta for every interval. Those exact grouped means drive the native chart below.

Exact notebook data

LapTime_Delta by tyre-life bin

HTML · CSS
-7.41s0–5
-3.52s5–10
-2.93s10–15
-2.69s15–20
-2.58s20–25
-2.18s25–30
-1.56s30–35
-1.44s35–40
-1.66s40–45
-1.87s45–50
Tyre life in laps
Exact grouped means printed by notebook cell 29
0-5 laps-7.41s5-10 laps-3.52s10-15 laps-2.93s15-20 laps-2.69s20-25 laps-2.58s25-30 laps-2.18s30-35 laps-1.56s35-40 laps-1.44s
Feature-engineering direction

The raw delta is relative to a reference pace. Rolling rate-of-change should be more useful than interpreting it as direct lap-over-lap tyre degradation.

08

TRAIN / TEST DRIFT

Only two features move by roughly five percent.

The notebook computes mean differences feature by feature. Every numeric feature remains under ten percent drift, and most stay below half a percent.

Exact notebook data

Train vs test mean drift

HTML · CSS
TyreLife
+0.02%
LapNumber
-0.24%
Position
-0.27%
LapTime (s)
+0.04%
LapTime_Delta
+5.10%
Cumulative_Degradation
-0.50%
RaceProgress
-0.29%
Position_Change
+5.18%
Stint
-0.27%
Exact percentage difference from notebook cell 32
FeatureMean difference
TyreLife+0.02%
LapNumber-0.24%
Position-0.27%
LapTime (s)+0.04%
LapTime_Delta+5.10%
Cumulative_Degradation-0.50%
RaceProgress-0.29%
Position_Change+5.18%
Stint-0.27%
Decision

No drift correction was required. Position_Change reaches 5.18% and LapTime_Delta 5.10%; all other mean differences are below 0.5%.

09

RACE-LEVEL PIT RATE

Circuit identity carries real strategy signal.

The notebook ranks every race by observed PitNextLap rate. Its printed top and bottom ten rows are recreated below, including Pre-Season Testing.

Exact notebook data

Pit rate by race

HTML · CSS
SELECTED CIRCUITChinese Grand Prix
Pit rate38.9%
Observed laps7,311
Approx. positives2,841

Highest rates

Lowest rates

Exact top and bottom ten rows printed by notebook cell 35
HighestChinese Grand Prix38.9%
HighestMonaco Grand Prix35.7%
HighestSpanish Grand Prix32.0%
LowestItalian Grand Prix13.2%
LowestMiami Grand Prix10.4%
LowestMexico City Grand Prix9.1%

Recommendation from the notebook: encode Race because circuit-level pit rates vary substantially. Remove Pre-Season Testing because its stops do not follow race-strategy logic.

10

BASELINE LIGHTGBM

EDA decisions flow into a time-aware baseline.

The notebook closes by encoding Driver, Compound, and Race; holding out 2025; training LightGBM; evaluating AUC; inspecting feature importance; and writing a submission.

01Encode

Driver · Compound · Race

02Train

346,246 rows · 2022-2024

03Validate

92,894 rows · 2025

04Score

ROC AUC · year-aware

0.89754Validation AUC
227Best iteration
0.284Validation positive rate
13Model features
Exact notebook data

LightGBM feature importance

HTML · CSS
RaceProgress801
Cumulative_Degradation801
LapTime (s)774
Race766
LapTime_Delta760
TyreLife553
LapNumber539
Driver381
Stint361
Position348
Exact top ten split-importance values from notebook cell 53
Submission rows188,165
Prediction mean0.3037
Prediction std0.3469
Prediction range0.0002-0.9931

NOTEBOOK CONCLUSION

The data is clean. The assumptions are the risky part.

HighRemove Pre-Season Testing rowsHighUse a year-based validation splitMediumEncode Compound and RaceMediumEngineer rolling degradation signalsLowSkip drift correction