Skip to content
TAXI NYC ’19, home

Optional

AI settings

Everything on this site works without AI. To try “Ask the data”, paste your own API key. Your browser sends it straight to the provider you pick. It is never sent to this site's server, never logged, and calls are billed to your account.

Provider
AI provider

Default. Cheapest and fastest; $1 / $5 per million input / output tokens.

No key saved for this provider. Create one at platform.claude.com. A key with a low spending limit is a good idea.

Method

Four rounds of cleaning, one regression

Everything on this site comes from re-running the 2021 notebook's rules, unchanged, on TLC's current copy of the 2019 data. This page lists every rule with the row counts side by side, then the model and how closely the revival reproduces it.

Cleaning

The funnel, rule by rule

Rows the notebook printed are shaded. The raw count differs by 0.24% because TLC's 2022 Parquet re-issue holds about 199,000 extra rows, almost all with missing values; after dropna() the two pipelines agree to within 0.003% at every checkpoint.

Steps run in notebook order; each “removed” bar is on a square-root scale so small rules stay visible. Model-stage rows differ from the analysis dataset because the shapefile join duplicates zones 56 and 103.

Cleaning funnel: rows left after each step, revived and notebook
#StepRemovedRows left (revived)NotebookDiff
1RawRaw 2019 yellow-taxi records84,598,44484,399,019+0.236%
2Round 1Drop rows with any missing value5,300,60179,297,84379,296,437+0.002%
3Round 1Trip distance must be positivetrip_distance > 0705,21878,592,625
4Round 1Passenger count between 1 and 6passenger_count BETWEEN 1 AND 61,431,24577,161,380
5Round 1Known rate code (drops 99)RatecodeID BETWEEN 1 AND 61,80377,159,577
6Round 1Fare must be positivefare_amount > 0156,42877,003,149
7Round 1Extra charge cannot be negativeextra >= 060077,002,549
8Round 1MTA tax of $0.50 to $1mta_tax BETWEEN 0.5 AND 1306,30776,696,242
9Round 1Tolls cannot be negativetolls_amount >= 076,696,242
10Round 1Improvement surcharge equals $0.30improvement_surcharge = 0.307,84876,688,394
11Round 1Total amount must be positivetotal_amount > 076,688,394
12Round 1Congestion surcharge cannot be negativecongestion_surcharge >= 076,688,394
13Round 1Pickup between 2018-01-01 and 2019-12-31pickup_dt BETWEEN TIMESTAMP '2018-01-01 00:00:00' AND TIMESTAMP '2019-12-31 23:59:59'96576,687,429
14Round 1Drop-off within 2019dropoff_dt BETWEEN TIMESTAMP '2019-01-01 00:00:00' AND TIMESTAMP '2019-12-31 23:59:59'1,15076,686,279
15Round 1Drop unknown vendor 4 (end of round 1)VendorID != 4197,47276,488,80776,487,438+0.002%
16Round 2Fare z-score at most 3fare <= mean + 3 sd = 296.1433 (notebook 296.1457)74776,488,06076,486,691+0.002%
17Round 2Drop exact duplicate rows376,488,05776,486,688+0.002%
18Round 3Fare at least the $2.50 flag-fallfare_amount >= 2.535176,487,70676,486,337+0.002%
19Round 3Drop-off after pick-uptravel_time > 03,18476,484,522
20Round 3At most 6 passengerspassenger_count <= 676,484,522
21Round 3Distance / minutes at most 50 (the 'mph' rule)trip_distance / travel_time <= 506,45476,478,06876,475,571+0.003%
22Round 3Trip at most 180 minutestravel_time <= 180206,18276,271,88676,269,392+0.003%
23Round 3Tip at most half the farefare_amount >= 2 * tip_amount596,04275,675,84475,673,363+0.003%
24MergeJoin taxi zones; drop trips from zone 264 to zone 26568175,675,163
25MergeDrop zones NA/NV; join weather (2019 pickups) and events489,46075,185,70375,183,226+0.003%
26MergeJoin collisions; drop Unknown boroughs (final analysis dataset)274,81474,910,88974,908,426+0.003%
27ModelJoin shapefile zone names (cell 255)74,942,128
28ModelDrop rows without an event count (dropna, cell 269)77374,941,355

Kept on purpose

Quirks of the 2021 rules

Try the rules

Would your trip survive?

The same rules, ported to TypeScript and unit-tested against rows printed in the notebook, run in your browser. Pick a preset or edit any field.

Clear a field to make it missing. Timestamps are New York wall time, “YYYY-MM-DD HH:MM:SS”.

Kept

This trip passes every cleaning rule of rounds 1 to 3.

travel_time
18.33 min
“mph” (mi/min)
0.13

Model

Elastic-net regression on 579 features

Spark MLlib's LinearRegression(maxIter=10, regParam=0.3, elasticNetParam=0.8) predicting time_duration_minutes, evaluated with a hand-written 10-fold cross-validation (cells 291-297). Features, in the order Spark assembled them:

Feature blocks of the model
BlockIndicesEncoding
numeric0–10Precipitation, snow, snow depth, TAVG, WT01, WT02, WT03, WT06, WT08, number_of_event, number_of_collision
weekday11–18Spark dayofweek of the pickup date (1 = Sunday), one-hot, size max + 1
hour19–42Pickup hour 0-23, one-hot
ratecode43–49RatecodeID 1-6, one-hot
passenger_count50–561-6, one-hot
pickup_zone57–314Shapefile zone name, StringIndexer by frequency, one-hot
vendor315–317VendorID 1-2, one-hot
dropoff_zone318–576Shapefile zone name, StringIndexer by frequency, one-hot
store_and_fwd_flag577–578N / Y via StringIndexer, one-hot
Training rows (revived)
74,941,355
Non-zero coefficients
83 of 579
2021 mean R² / RMSE
0.3665 / 9.175
2021 coefficients on revived folds
0.3681 / 9.179
2026 refit, same objective
0.3680 / 9.180
Zone coefficients, 2021 vs refit
r = 0.99996

Scoring the 2021 coefficients on the revived folds only works if the cleaning, the feature order and the frequency-ordered zone index all match the notebook; they do, to the third decimal of R². The refit solves Spark's standardised objective to convergence with coordinate descent on exact X'X sums (scripts/fit_model.py) and lands on the same model. Folds are a hash of each trip, because Spark's randomSplit cannot be replayed.

R² by fold

0.36460.36720.369712345678910
2021 notebook2021 coefficients, revived data2026 refitx axis: fold

RMSE by fold (minutes)

9.1619.1809.20012345678910
2021 notebook2021 coefficients, revived data2026 refitx axis: fold

Hindsight

The penalty did most of the talking

Same features and elastic-net mix, different regParam. The 2021 choice of 0.3 kept 83 coefficients; a penalty thirty times lighter keeps most zones and lifts cross-validated R² from 0.37 to about 0.45. The revival reports this rather than changing the model.

Regularisation path: mean 10-fold R², RMSE and non-zero coefficients
regParamCV R²CV RMSENon-zero coefficients
0.010.44528.601502
0.030.44318.617423
0.10.42658.745277
0.32021 choice0.36809.18083
10.30789.6077
30.188110.4054