Method
Four rounds of cleaning, one regression
Everything on this site comes from re-running the 2021 notebook's rules, unchanged, on TLC's current copy of the 2019 data. This page lists every rule with the row counts side by side, then the model and how closely the revival reproduces it.
Cleaning
The funnel, rule by rule
Rows the notebook printed are shaded. The raw count differs by 0.24% because TLC's 2022 Parquet re-issue holds about 199,000 extra rows, almost all with missing values; after dropna() the two pipelines agree to within 0.003% at every checkpoint.
Steps run in notebook order; each “removed” bar is on a square-root scale so small rules stay visible. Model-stage rows differ from the analysis dataset because the shapefile join duplicates zones 56 and 103.
| # | Step | Removed | Rows left (revived) | Notebook | Diff |
|---|---|---|---|---|---|
| 1 | RawRaw 2019 yellow-taxi records | 84,598,444 | 84,399,019 | +0.236% | |
| 2 | Round 1Drop rows with any missing value | 5,300,601 | 79,297,843 | 79,296,437 | +0.002% |
| 3 | Round 1Trip distance must be positivetrip_distance > 0 | 705,218 | 78,592,625 | ||
| 4 | Round 1Passenger count between 1 and 6passenger_count BETWEEN 1 AND 6 | 1,431,245 | 77,161,380 | ||
| 5 | Round 1Known rate code (drops 99)RatecodeID BETWEEN 1 AND 6 | 1,803 | 77,159,577 | ||
| 6 | Round 1Fare must be positivefare_amount > 0 | 156,428 | 77,003,149 | ||
| 7 | Round 1Extra charge cannot be negativeextra >= 0 | 600 | 77,002,549 | ||
| 8 | Round 1MTA tax of $0.50 to $1mta_tax BETWEEN 0.5 AND 1 | 306,307 | 76,696,242 | ||
| 9 | Round 1Tolls cannot be negativetolls_amount >= 0 | 76,696,242 | |||
| 10 | Round 1Improvement surcharge equals $0.30improvement_surcharge = 0.30 | 7,848 | 76,688,394 | ||
| 11 | Round 1Total amount must be positivetotal_amount > 0 | 76,688,394 | |||
| 12 | Round 1Congestion surcharge cannot be negativecongestion_surcharge >= 0 | 76,688,394 | |||
| 13 | Round 1Pickup between 2018-01-01 and 2019-12-31pickup_dt BETWEEN TIMESTAMP '2018-01-01 00:00:00' AND TIMESTAMP '2019-12-31 23:59:59' | 965 | 76,687,429 | ||
| 14 | Round 1Drop-off within 2019dropoff_dt BETWEEN TIMESTAMP '2019-01-01 00:00:00' AND TIMESTAMP '2019-12-31 23:59:59' | 1,150 | 76,686,279 | ||
| 15 | Round 1Drop unknown vendor 4 (end of round 1)VendorID != 4 | 197,472 | 76,488,807 | 76,487,438 | +0.002% |
| 16 | Round 2Fare z-score at most 3fare <= mean + 3 sd = 296.1433 (notebook 296.1457) | 747 | 76,488,060 | 76,486,691 | +0.002% |
| 17 | Round 2Drop exact duplicate rows | 3 | 76,488,057 | 76,486,688 | +0.002% |
| 18 | Round 3Fare at least the $2.50 flag-fallfare_amount >= 2.5 | 351 | 76,487,706 | 76,486,337 | +0.002% |
| 19 | Round 3Drop-off after pick-uptravel_time > 0 | 3,184 | 76,484,522 | ||
| 20 | Round 3At most 6 passengerspassenger_count <= 6 | 76,484,522 | |||
| 21 | Round 3Distance / minutes at most 50 (the 'mph' rule)trip_distance / travel_time <= 50 | 6,454 | 76,478,068 | 76,475,571 | +0.003% |
| 22 | Round 3Trip at most 180 minutestravel_time <= 180 | 206,182 | 76,271,886 | 76,269,392 | +0.003% |
| 23 | Round 3Tip at most half the farefare_amount >= 2 * tip_amount | 596,042 | 75,675,844 | 75,673,363 | +0.003% |
| 24 | MergeJoin taxi zones; drop trips from zone 264 to zone 265 | 681 | 75,675,163 | ||
| 25 | MergeDrop zones NA/NV; join weather (2019 pickups) and events | 489,460 | 75,185,703 | 75,183,226 | +0.003% |
| 26 | MergeJoin collisions; drop Unknown boroughs (final analysis dataset) | 274,814 | 74,910,889 | 74,908,426 | +0.003% |
| 27 | ModelJoin shapefile zone names (cell 255) | 74,942,128 | |||
| 28 | ModelDrop rows without an event count (dropna, cell 269) | 773 | 74,941,355 |
Kept on purpose
Quirks of the 2021 rules
Try the rules
Would your trip survive?
The same rules, ported to TypeScript and unit-tested against rows printed in the notebook, run in your browser. Pick a preset or edit any field.
Clear a field to make it missing. Timestamps are New York wall time, “YYYY-MM-DD HH:MM:SS”.
Kept
This trip passes every cleaning rule of rounds 1 to 3.
- travel_time
- 18.33 min
- “mph” (mi/min)
- 0.13
Model
Elastic-net regression on 579 features
Spark MLlib's LinearRegression(maxIter=10, regParam=0.3, elasticNetParam=0.8) predicting time_duration_minutes, evaluated with a hand-written 10-fold cross-validation (cells 291-297). Features, in the order Spark assembled them:
| Block | Indices | Encoding |
|---|---|---|
| numeric | 0–10 | Precipitation, snow, snow depth, TAVG, WT01, WT02, WT03, WT06, WT08, number_of_event, number_of_collision |
| weekday | 11–18 | Spark dayofweek of the pickup date (1 = Sunday), one-hot, size max + 1 |
| hour | 19–42 | Pickup hour 0-23, one-hot |
| ratecode | 43–49 | RatecodeID 1-6, one-hot |
| passenger_count | 50–56 | 1-6, one-hot |
| pickup_zone | 57–314 | Shapefile zone name, StringIndexer by frequency, one-hot |
| vendor | 315–317 | VendorID 1-2, one-hot |
| dropoff_zone | 318–576 | Shapefile zone name, StringIndexer by frequency, one-hot |
| store_and_fwd_flag | 577–578 | N / Y via StringIndexer, one-hot |
- Training rows (revived)
- 74,941,355
- Non-zero coefficients
- 83 of 579
- 2021 mean R² / RMSE
- 0.3665 / 9.175
- 2021 coefficients on revived folds
- 0.3681 / 9.179
- 2026 refit, same objective
- 0.3680 / 9.180
- Zone coefficients, 2021 vs refit
- r = 0.99996
Scoring the 2021 coefficients on the revived folds only works if the cleaning, the feature order and the frequency-ordered zone index all match the notebook; they do, to the third decimal of R². The refit solves Spark's standardised objective to convergence with coordinate descent on exact X'X sums (scripts/fit_model.py) and lands on the same model. Folds are a hash of each trip, because Spark's randomSplit cannot be replayed.
R² by fold
RMSE by fold (minutes)
Hindsight
The penalty did most of the talking
Same features and elastic-net mix, different regParam. The 2021 choice of 0.3 kept 83 coefficients; a penalty thirty times lighter keeps most zones and lifts cross-validated R² from 0.37 to about 0.45. The revival reports this rather than changing the model.
| regParam | CV R² | CV RMSE | Non-zero coefficients |
|---|---|---|---|
| 0.01 | 0.4452 | 8.601 | 502 |
| 0.03 | 0.4431 | 8.617 | 423 |
| 0.1 | 0.4265 | 8.745 | 277 |
| 0.32021 choice | 0.3680 | 9.180 | 83 |
| 1 | 0.3078 | 9.607 | 7 |
| 3 | 0.1881 | 10.405 | 4 |