Evaluation
How good is the model, honestly?
The 2021 notebook judged its regression by 10-fold cross-validation on a random split. This page asks harder questions: how it does on months it never saw, how sure we can be of each coefficient, where its errors pile up, and how wide an honest prediction interval has to be. Every interval states its method, sample size and seed.
Temporal hold-out
Train on January–October, test on November–December
Models fitted on 55,804,179 trips from January to October (folds 1–9), scored on all 12,936,277 trips of November and December. Intervals come from a cluster bootstrap that resamples whole days (61 test days, B = 2,000, seed 20190101): trips on the same day share weather and traffic, so they are not independent.
| Model (Nov–Dec) | RMSE (95% CI) | MAE (95% CI) | R² (95% CI) | ΔRMSE vs 2021 spec (paired) |
|---|---|---|---|---|
| Route × hour median | 6.24 (6.00 to 6.47) | 3.95 (3.83 to 4.07) | 0.719 (0.707 to 0.732) | -3.23 (-3.34 to -3.10) |
| Same features, lighter penalty | 8.95 (8.68 to 9.20) | 6.33 (6.17 to 6.48) | 0.422 (0.408 to 0.435) | -0.52 (-0.56 to -0.48) |
| 2021 coefficients | 9.43 (9.15 to 9.70) | 6.80 (6.63 to 6.96) | 0.358 (0.345 to 0.370) | -0.03 (-0.05 to -0.02) |
| 2021 specification | 9.47 (9.18 to 9.73) | 6.79 (6.61 to 6.95) | 0.353 (0.341 to 0.365) | reference |
| Mean only | 11.78 (11.40 to 12.13) | 8.24 (8.09 to 8.38) | -0.001 (-0.003 to 0.000) | 2.31 (2.18 to 2.44) |
In-period random hold-out (Jan–Oct fold 0)
6,200,899 trips over 288 days, the same bootstrap. The 2021 coefficients were trained on a random 90% of all of 2019, so on any 2019 test set they are not a true hold-out. The gap to the Jan–Oct refit is within noise.
- Route × hour median: RMSE 5.77 (5.67 to 5.88), R² 0.747 (0.742 to 0.753)
- Same features, lighter penalty: RMSE 8.53 (8.43 to 8.64), R² 0.448 (0.442 to 0.453)
- 2021 coefficients: RMSE 9.12 (9.01 to 9.24), R² 0.369 (0.364 to 0.374)
- 2021 specification: RMSE 9.12 (9.01 to 9.24), R² 0.368 (0.363 to 0.374)
- Mean only: RMSE 11.48 (11.33 to 11.64), R² 0.000 (0.000 to 0.000)
Inference
Coefficients with honest uncertainty
Penalised coefficients have no standard errors worth quoting, so this is the unpenalised OLS counterpart: the same 579-column design on all 74,941,355 model rows, one reference level dropped per block (567 parameters, in-sample R² 0.446). Three standard errors per coefficient: classical, heteroskedasticity-robust HC3, and clustered by pickup day (351 days, t critical value 1.967). The numpy code was checked against statsmodels before it ran (scripts/sparse_ols.py).
| Feature | Estimate | SE classical | SE HC3 | SE by day | Day / HC3 | 95% CI, by day |
|---|---|---|---|---|---|---|
| Precipitation | 0.400 | 0.0039 | 0.0040 | 0.191 | 47× | 0.024 to 0.776 |
| Snow | −0.316 | 0.0046 | 0.0043 | 0.131 | 30× | −0.574 to −0.058 |
| Snow depth | −0.035 | 0.0039 | 0.0037 | 0.112 | 30× | −0.255 to 0.185 |
| TAVG | −0.0032 | 0.00010 | 0.00010 | 0.0053 | 52× | −0.014 to 0.0072 |
| WT01 | 0.045 | 0.0026 | 0.0026 | 0.112 | 43× | −0.175 to 0.264 |
| WT02 | −0.132 | 0.0073 | 0.0072 | 0.238 | 33× | −0.600 to 0.336 |
| WT03 | −0.292 | 0.0041 | 0.0041 | 0.143 | 35× | −0.574 to −0.0096 |
| WT06 | −0.543 | 0.0080 | 0.0077 | 0.366 | 48× | −1.26 to 0.176 |
| WT08 | −0.0095 | 0.0030 | 0.0030 | 0.122 | 41× | −0.249 to 0.230 |
| number_of_event | 0.00035 | 0.00001 | 0.00001 | 0.00067 | 48× | −0.00097 to 0.0017 |
| number_of_collision | 0.031 | 0.00007 | 0.00007 | 0.0032 | 48× | 0.025 to 0.037 |
Largest zone effects (zones with 1,000+ trips)
| Murray Hill-Queens | 15.90 | 14.39 to 17.41 |
|---|---|---|
| LaGuardia Airport | 15.33 | 15.31 to 15.35 |
| Queens Village | 12.57 | 11.53 to 13.61 |
| Saint Albans | 12.14 | 11.44 to 12.83 |
| Rosedale | 11.57 | 10.47 to 12.67 |
| Maspeth | −9.08 | −9.46 to −8.70 |
| Gowanus | −9.56 | −9.88 to −9.23 |
| Astoria Park | −10.49 | −11.36 to −9.62 |
| New Dorp/Midland Beach | 34.29 | 33.49 to 35.09 |
|---|---|---|
| Saint George/New Brighton | 33.81 | 32.94 to 34.68 |
| Heartland Village/Todt Hill | 30.85 | 29.99 to 31.71 |
| Bloomfield/Emerson Hill | 29.10 | 28.13 to 30.07 |
| Westerleigh | 28.34 | 27.46 to 29.23 |
| Saint Michaels Cemetery/Woodside | −3.32 | −4.15 to −2.50 |
| Springfield Gardens South | −4.30 | −4.48 to −4.12 |
| Baisley Park | −5.34 | −5.47 to −5.21 |
All 567 coefficients with every standard error: CSV or browse.
Diagnostics
Where the errors pile up
Residuals of every 2019 model row, for the 2021 coefficients and for the OLS fit. A well-specified linear model would show a flat band around zero of constant width. These show a skewed, fanning band: long trips are under-predicted, and outer-borough pickups are far noisier than Manhattan ones.
74,941,355 trips · residual SD 9.18 min · MAE 6.70 min
Residuals against fitted values
Normal QQ plot
Residual SD by pickup borough
- Manhattan8.5 minn 68,963,489
- Brooklyn13.6 minn 809,938
- Queens15.3 minn 5,067,090
- Bronx16.9 minn 98,435
- Staten Island30.7 minn 2,403
Residual SD by pickup hour (minutes)
Prediction intervals
Intervals that hold their coverage
Split-conformal prediction wraps any model: hold back calibration trips, look at their residuals, and use the right order statistic as the interval. The guarantee is about coverage on average, so it says nothing about whether short and long trips are each covered. Conditioning on a group (Mondrian conformal) fixes that for the groups you choose. Scheme 1 calibrates the 2021 coefficients on 7,493,562 random 2019 trips and tests on 7,497,013 others. Scheme 2 refits on Jan–Oct and tests on 12,936,277 Nov–Dec trips. See DR-003.
| Method | 80% target | 90% target | 95% target |
|---|---|---|---|
| Scheme 1: 2021 coefficients, random 2019 test fold | |||
| Global, symmetric | 80.0% (79.6% to 80.4%) · 18.7 min | 90.0% (89.7% to 90.2%) · 25.0 min | 95.0% (94.8% to 95.2%) · 31.2 min |
| Mondrian by predicted decile | 80.0% (79.8% to 80.2%) · 20.3 min | 90.0% (89.9% to 90.1%) · 26.6 min | 95.0% (94.9% to 95.1%) · 32.5 min |
| Mondrian by borough × decile | 80.0% (79.8% to 80.2%) · 20.2 min | 90.0% (89.9% to 90.1%) · 26.4 min | 95.0% (94.9% to 95.1%) · 32.5 min |
| Scheme 2: Jan–Oct refit, tested on Nov–Dec | |||
| Global, symmetric | 79.3% (78.4% to 80.2%) · 18.6 min | 88.5% (87.9% to 89.1%) · 24.6 min | 94.1% (93.7% to 94.6%) · 30.5 min |
| Mondrian by predicted decile | 78.4% (78.0% to 78.9%) · 19.4 min | 88.9% (88.5% to 89.3%) · 25.4 min | 94.3% (94.1% to 94.6%) · 31.1 min |
| Mondrian by borough × decile | 78.3% (77.9% to 78.9%) · 19.2 min | 88.8% (88.4% to 89.2%) · 25.2 min | 94.2% (94.0% to 94.5%) · 31.0 min |