Records
The analytics database
Every chart on this site reads from one read-only SQLite file, web/data/analytics.db, built by scripts/build_analytics.py from the cleaned trips, with the evidence and data-quality tables added by scripts/build_evidence_tables.py. It holds 37 tables and 312,194 rows of aggregates; no individual trip is stored.
- CSV
Borough flows
borough_flows · 36 rows · 4 columns
Trips between boroughs (the 2021 'route' feature), with median minutes.
pickup_borough · dropoff_borough · trips · median_min
- CSV
Cleaning funnel
cleaning_funnel · 28 rows · 7 columns
Rows left after every cleaning rule, revived pipeline next to the counts printed by the 2021 notebook.
step · stage · label · rule · revived_rows · notebook_rows · difference_pct
- CSV
Collisions
collisions_hourly · 32,705 rows · 4 columns
NYPD motor-vehicle collisions per borough, day and hour (2021 BigQuery export).
borough · date · hour · collisions
- CSV
Conformal interval table
conformal_bins · 312 rows · 10 columns
Split-conformal offsets per calibration bin: global symmetric, Mondrian by predicted decile, and Mondrian by pickup borough x predicted bin, at 80%, 90% and 95%.
scheme · method · level · borough · bin · pred_lo · pred_hi · n_cal · q_lo · q_hi
- CSV
Conformal coverage
conformal_coverage · 1,440 rows · 11 columns
Empirical coverage of each conformal method on held-out trips, overall and by bin, borough and hour, with mean interval width (lower end clipped at 0) and a 95% interval from a bootstrap that resamples whole test days (B = 2,000, seed 20190101).
scheme · method · level · group_type · group_value · trips · covered · mean_width · days · ci_low · ci_high
- CSV
Conformal coverage per test day
conformal_coverage_daily · 3,672 rows · 6 columns
Trips and covered trips per test day for each scheme, method and level: the units the coverage bootstrap resamples.
scheme · method · level · date · trips · covered
- CSV
Daily series
daily · 365 rows · 20 columns
One row per day of 2019: trips, duration statistics, NOAA Central Park weather, permitted events and collisions (summed over boroughs).
date · isodow · trips · median_min · mean_min · mean_miles · mean_fare · precipitation · snow · snow_depth · tavg · tmax · tmin · wt01 · wt02 · wt03 · wt06 · wt08 · events · collisions
- CSV
Daily x borough
daily_borough · 1,825 rows · 6 columns
Pickups, median minutes, permitted events and collisions per day and borough.
date · borough · pickups · median_min · events · collisions
- CSV
January missing surcharge
dq_january · 31 rows · 3 columns
Rows and missing congestion surcharges per pickup day in the January 2019 file.
date · rows · missing_congestion
- CSV
Missing values by source file
dq_missing · 60 rows · 4 columns
Missing values per column in each monthly TLC file of 2019.
source_file · column_name · missing · rows
- CSV
Residual plausibility checks
dq_residual_checks · 8 rows · 10 columns
Implausible records the 2021 rules let through, counted on the final analysis dataset by vendor (reported, not removed).
key · label · condition · reason · rows · share · vendor1_rows · vendor2_rows · vendor1_trips · vendor2_trips
- CSV
Offending values per rule
dq_rule_values · 87 rows · 4 columns
The most common values each cleaning rule removes.
key · rank · value · rows
- CSV
Data-quality rules
dq_rules · 21 rows · 11 columns
Every 2021 cleaning rule with its reason, rows removed in sequence, rows failing it on its own and rows failing only it.
step · key · stage · label · rule · reason · applied_to · applied_rows · removed_in_sequence · fails_alone · fails_only_this
- CSV
Congestion surcharge convention
dq_surcharge_by_month · 24 rows · 4 columns
Per month and vendor: trips whose total leaves out the recorded congestion surcharge.
month · vendor · rows · trips
- CSV
Borough-day duration index
effects_borough_daily · 2,001 rows · 12 columns
Per pickup borough and day: trips, the duration index, permitted events and collisions in that borough.
date · borough · trips · indexed_trips · mean_min · median_min · mean_log_ratio · sd_log_ratio · precipitation · tavg · events · collisions
- CSV
Daily duration index
effects_daily · 351 rows · 13 columns
Per day: trips, mean minutes and the composition-adjusted duration index (mean log of minutes / route-hour median), with weather, events and collisions.
date · trips · indexed_trips · mean_min · median_min · mean_log_ratio · sd_log_ratio · precipitation · snow · snow_depth · tavg · events · collisions
- CSV
Permitted events
events_daily · 1,823 rows · 3 columns
Permitted events starting each day per borough (NYC Open Data bkfu-528j, current version).
borough · date · number_of_event
- CSV
Evidence metadata
evidence_meta · 8 rows · 2 columns
Key-value summaries (JSON) of the regression, hold-out, effects and data-quality analyses: sample sizes, references, splits and self-checks.
key · value
- CSV
Hold-out errors per day
holdout_daily · 1,745 rows · 9 columns
Per test day and model: trips and sums of errors, squared errors, absolute errors, y and y squared. The website bootstraps whole days from these.
split · model · date · trips · sum_err · sum_sq_err · sum_abs_err · sum_y · sum_y2
- CSV
Hold-out models
holdout_models · 5 rows · 2 columns
The models compared on the temporal (Nov-Dec) and in-period (Jan-Oct fold 0) hold-outs.
model · label
- CSV
Table descriptions
meta_tables · 36 rows · 3 columns
The descriptions shown on this page.
name · title · description
- CSV
Model coefficients
model_coefficients · 579 rows · 6 columns
All 579 regression coefficients in VectorAssembler order: 2021 fold 1 next to the 2026 refit.
feature_index · block · level · label · original_2021 · refit_2026
- CSV
Cross-validation folds
model_folds · 10 rows · 9 columns
R^2 and RMSE per fold: the 2021 notebook, the 2021 coefficients scored on the revived folds, and the 2026 refit.
fold · test_rows · notebook_r2 · notebook_rmse · original_on_revived_r2 · original_on_revived_rmse · refit_r2 · refit_rmse · refit_nonzero
- CSV
Regularisation path
model_path · 6 rows · 5 columns
Mean 10-fold R^2, RMSE and number of non-zero coefficients for several regParam values at the notebook's elasticNetParam = 0.8.
reg_param · elastic_net_param · cv_r2 · cv_rmse · nonzero
- CSV
Zone name index
model_zone_index · 517 rows · 4 columns
Spark StringIndexer order of pickup and drop-off zone names (most trips first), rebuilt from the revived data.
side · idx · zone · trips
- CSV
OLS coefficients with robust SEs
ols_coefficients · 567 rows · 13 columns
Unpenalised OLS counterpart of the 2021 model on all 2019 model rows: estimates with classical, HC3 and day-clustered standard errors and 95% intervals. One reference level per block is omitted.
feature_index · block · level · label · trips · estimate · se_classical · se_hc3 · se_cluster_day · ci_low_hc3 · ci_high_hc3 · ci_low_cluster · ci_high_cluster
- CSV
Residuals by fitted value
residual_bins · 126 rows · 8 columns
Mean and 10th/50th/90th percentile residual per 1-minute fitted bin (bins with at least 200 trips).
model · fitted_lo · fitted_hi · trips · mean_resid · p10 · p50 · p90
- CSV
Residual spread by group
residual_groups · 298 rows · 7 columns
Trips, mean residual, residual SD and MAE per pickup borough, pickup hour and borough x hour.
model · group_type · group_value · trips · mean_resid · sd_resid · mae
- CSV
Residuals vs fitted (2D histogram)
residual_hist2d · 9,133 rows · 4 columns
Trips per 1-minute fitted bin and 2-minute residual bin, for the 2021 coefficients and the OLS fit.
model · fitted_lo · resid_lo · trips
- CSV
Residual quantiles (QQ)
residual_qq · 94 rows · 4 columns
Residual quantiles against the normal quantiles with the same mean and standard deviation.
model · p · sample_q · normal_q
- CSV
Route x hour
route_hourly · 107,915 rows · 5 columns
Hourly trips and median minutes for zone pairs with at least 1,000 trips in 2019.
pu_id · do_id · hour · trips · median_min
- CSV
Zone-to-zone routes
routes · 45,836 rows · 8 columns
Every pickup zone to drop-off zone pair seen in 2019 with trips, median and mean minutes, distance, fare and speed.
pu_id · do_id · trips · median_min · mean_min · mean_miles · mean_fare · mean_mph
- CSV
NOAA weather
weather · 365 rows · 12 columns
Daily Central Park weather after the notebook's cleaning: TAVG = (TMAX + TMIN) / 2, gaps filled with 0, WT04 dropped.
date · precipitation · snow · snow_depth · tavg · tmax · tmin · wt01 · wt02 · wt03 · wt06 · wt08
- CSV
Weekday x hour x vendor
weekday_hour_vendor · 504 rows · 6 columns
Mean and median trip minutes per ISO weekday, pickup hour and vendor (0 = both), as in the 2021 vendor line charts.
vendor · isodow · hour · trips · mean_min · median_min
- CSV
Zone x weekday x hour
zone_hourly · 98,357 rows · 7 columns
Trips and trip-duration statistics per zone, ISO weekday (1 = Monday, 0 = all days) and hour (24 = all day), for pickups (pickup time) and drop-offs (drop-off time).
side · location_id · isodow · hour · trips · median_min · mean_min
- CSV
Zone x vendor
zone_vendor · 1,039 rows · 4 columns
Yearly trips per zone and vendor, the data behind the 2021 vendor choropleths.
side · location_id · vendor · trips
- CSV
Taxi zones
zones · 265 rows · 10 columns
One row per TLC taxi zone: borough, service zone, the zone name the 2021 model used, map centroid and yearly pickups/drop-offs.
location_id · borough · zone · service_zone · model_zone_name · has_polygon · centroid_lon · centroid_lat · pickups · dropoffs