Skip to content
TAXI NYC ’19, home

Optional

AI settings

Everything on this site works without AI. To try “Ask the data”, paste your own API key. Your browser sends it straight to the provider you pick. It is never sent to this site's server, never logged, and calls are billed to your account.

Provider
AI provider

Default. Cheapest and fastest; $1 / $5 per million input / output tokens.

No key saved for this provider. Create one at platform.claude.com. A key with a low spending limit is a good idea.

Evaluation

How good is the model, honestly?

The 2021 notebook judged its regression by 10-fold cross-validation on a random split. This page asks harder questions: how it does on months it never saw, how sure we can be of each coefficient, where its errors pile up, and how wide an honest prediction interval has to be. Every interval states its method, sample size and seed.

Temporal hold-out

Train on January–October, test on November–December

Models fitted on 55,804,179 trips from January to October (folds 1–9), scored on all 12,936,277 trips of November and December. Intervals come from a cluster bootstrap that resamples whole days (61 test days, B = 2,000, seed 20190101): trips on the same day share weather and traffic, so they are not independent.

Nov–Dec (temporal)Jan–Oct fold 0 (random)RMSE in minutes, lower is better. Right-hand figures show Nov–Dec with its 95% interval
Temporal hold-out metrics with 95% day-bootstrap intervals
Model (Nov–Dec)RMSE (95% CI)MAE (95% CI)R² (95% CI)ΔRMSE vs 2021 spec (paired)
Route × hour median6.24 (6.00 to 6.47)3.95 (3.83 to 4.07)0.719 (0.707 to 0.732)-3.23 (-3.34 to -3.10)
Same features, lighter penalty8.95 (8.68 to 9.20)6.33 (6.17 to 6.48)0.422 (0.408 to 0.435)-0.52 (-0.56 to -0.48)
2021 coefficients9.43 (9.15 to 9.70)6.80 (6.63 to 6.96)0.358 (0.345 to 0.370)-0.03 (-0.05 to -0.02)
2021 specification9.47 (9.18 to 9.73)6.79 (6.61 to 6.95)0.353 (0.341 to 0.365)reference
Mean only11.78 (11.40 to 12.13)8.24 (8.09 to 8.38)-0.001 (-0.003 to 0.000)2.31 (2.18 to 2.44)
In-period random hold-out (Jan–Oct fold 0)

6,200,899 trips over 288 days, the same bootstrap. The 2021 coefficients were trained on a random 90% of all of 2019, so on any 2019 test set they are not a true hold-out. The gap to the Jan–Oct refit is within noise.

  • Route × hour median: RMSE 5.77 (5.67 to 5.88), R² 0.747 (0.742 to 0.753)
  • Same features, lighter penalty: RMSE 8.53 (8.43 to 8.64), R² 0.448 (0.442 to 0.453)
  • 2021 coefficients: RMSE 9.12 (9.01 to 9.24), R² 0.369 (0.364 to 0.374)
  • 2021 specification: RMSE 9.12 (9.01 to 9.24), R² 0.368 (0.363 to 0.374)
  • Mean only: RMSE 11.48 (11.33 to 11.64), R² 0.000 (0.000 to 0.000)

Inference

Coefficients with honest uncertainty

Penalised coefficients have no standard errors worth quoting, so this is the unpenalised OLS counterpart: the same 579-column design on all 74,941,355 model rows, one reference level dropped per block (567 parameters, in-sample R² 0.446). Three standard errors per coefficient: classical, heteroskedasticity-robust HC3, and clustered by pickup day (351 days, t critical value 1.967). The numpy code was checked against statsmodels before it ran (scripts/sparse_ols.py).

Weather, events and collisions: minutes added per unit. Every trip on a day shares these values.
FeatureEstimateSE classicalSE HC3SE by dayDay / HC395% CI, by day
Precipitation0.4000.00390.00400.19147×0.024 to 0.776
Snow−0.3160.00460.00430.13130×−0.574 to −0.058
Snow depth−0.0350.00390.00370.11230×−0.255 to 0.185
TAVG−0.00320.000100.000100.005352×−0.014 to 0.0072
WT010.0450.00260.00260.11243×−0.175 to 0.264
WT02−0.1320.00730.00720.23833×−0.600 to 0.336
WT03−0.2920.00410.00410.14335×−0.574 to −0.0096
WT06−0.5430.00800.00770.36648×−1.26 to 0.176
WT08−0.00950.00300.00300.12241×−0.249 to 0.230
number_of_event0.000350.000010.000010.0006748×−0.00097 to 0.0017
number_of_collision0.0310.000070.000070.003248×0.025 to 0.037
Minutes relative to 18:00, day-clustered 95% CIVertical line: the reference hour, 18:00
Minutes relative to Thursday, day-clustered 95% CIVertical line: the reference day, Thursday

Largest zone effects (zones with 1,000+ trips)

Pickup zone, minutes relative to Upper East Side South (HC3 95% CI)
Murray Hill-Queens15.9014.39 to 17.41
LaGuardia Airport15.3315.31 to 15.35
Queens Village12.5711.53 to 13.61
Saint Albans12.1411.44 to 12.83
Rosedale11.5710.47 to 12.67
Maspeth−9.08−9.46 to −8.70
Gowanus−9.56−9.88 to −9.23
Astoria Park−10.49−11.36 to −9.62
Drop-off zone, minutes relative to Upper East Side North (HC3 95% CI)
New Dorp/Midland Beach34.2933.49 to 35.09
Saint George/New Brighton33.8132.94 to 34.68
Heartland Village/Todt Hill30.8529.99 to 31.71
Bloomfield/Emerson Hill29.1028.13 to 30.07
Westerleigh28.3427.46 to 29.23
Saint Michaels Cemetery/Woodside−3.32−4.15 to −2.50
Springfield Gardens South−4.30−4.48 to −4.12
Baisley Park−5.34−5.47 to −5.21

All 567 coefficients with every standard error: CSV or browse.

Diagnostics

Where the errors pile up

Residuals of every 2019 model row, for the 2021 coefficients and for the OLS fit. A well-specified linear model would show a flat band around zero of constant width. These show a skewed, fanning band: long trips are under-predicted, and outer-borough pickups are far noisier than Manhattan ones.

Model

74,941,355 trips · residual SD 9.18 min · MAE 6.70 min

Residuals against fitted values

-40040800204060Fitted minutesResidual, minutes
Shading shows trips per cell on a log scale, with the strongest colour for the densest cells. The solid line is the median residual in each 1-minute band of fitted values with at least 20,000 trips, the dashed lines the 10th and 90th percentiles, and the red line the mean residual. Trips outside the plotted window are counted but not drawn.

Normal QQ plot

-40-40-20-2000202040406060Normal quantile (same mean and SD)Residual quantile, minutes
47 quantiles from 0.1% to 99.9%. Points above the dashed line on the right are a long right tail: some rides take far longer than a normal error would allow.

Residual SD by pickup borough

  • Manhattan8.5 minn 68,963,489
  • Brooklyn13.6 minn 809,938
  • Queens15.3 minn 5,067,090
  • Bronx16.9 minn 98,435
  • Staten Island30.7 minn 2,403
Mean residual by borough: Manhattan −0.1, Brooklyn +0.3, Queens +1.1, Bronx +8.6, Staten Island +29.6 minutes.

Residual SD by pickup hour (minutes)

06120006121823
Borough × hour explains 4.4% of the variation in squared residuals (Breusch–Pagan-style LM = 3,260,599 on 119 df). At this sample size any test rejects constant variance, so the effect size is the useful number.

Prediction intervals

Intervals that hold their coverage

Split-conformal prediction wraps any model: hold back calibration trips, look at their residuals, and use the right order statistic as the interval. The guarantee is about coverage on average, so it says nothing about whether short and long trips are each covered. Conditioning on a group (Mondrian conformal) fixes that for the groups you choose. Scheme 1 calibrates the 2021 coefficients on 7,493,562 random 2019 trips and tests on 7,497,013 others. Scheme 2 refits on Jan–Oct and tests on 12,936,277 Nov–Dec trips. See DR-003.

Empirical coverage with a 95% interval that resamples whole test days (347 days in scheme 1, 61 in scheme 2, B = 2,000, seed 20190101), and mean width in minutes (lower end clipped at 0).
Method80% target90% target95% target
Scheme 1: 2021 coefficients, random 2019 test fold
Global, symmetric80.0% (79.6% to 80.4%) · 18.7 min90.0% (89.7% to 90.2%) · 25.0 min95.0% (94.8% to 95.2%) · 31.2 min
Mondrian by predicted decile80.0% (79.8% to 80.2%) · 20.3 min90.0% (89.9% to 90.1%) · 26.6 min95.0% (94.9% to 95.1%) · 32.5 min
Mondrian by borough × decile80.0% (79.8% to 80.2%) · 20.2 min90.0% (89.9% to 90.1%) · 26.4 min95.0% (94.9% to 95.1%) · 32.5 min
Scheme 2: Jan–Oct refit, tested on Nov–Dec
Global, symmetric79.3% (78.4% to 80.2%) · 18.6 min88.5% (87.9% to 89.1%) · 24.6 min94.1% (93.7% to 94.6%) · 30.5 min
Mondrian by predicted decile78.4% (78.0% to 78.9%) · 19.4 min88.9% (88.5% to 89.3%) · 25.4 min94.3% (94.1% to 94.6%) · 31.1 min
Mondrian by borough × decile78.3% (77.9% to 78.9%) · 19.2 min88.8% (88.4% to 89.2%) · 25.2 min94.2% (94.0% to 94.5%) · 31.0 min
Global, symmetricMondrian by predicted decileMondrian by borough × decileCoverage of 90% intervals by predicted-duration decile (scheme 1)Vertical line: 90% target
Global, symmetricMondrian by predicted decileMondrian by borough × decileCoverage of 90% intervals by pickup borough (scheme 1)Vertical line: 90% target