Skip to content
TAXI NYC ’19, home

Optional

AI settings

Everything on this site works without AI. To try “Ask the data”, paste your own API key. Your browser sends it straight to the provider you pick. It is never sent to this site's server, never logged, and calls are billed to your account.

Provider
AI provider

Default. Cheapest and fastest; $1 / $5 per million input / output tokens.

No key saved for this provider. Create one at platform.claude.com. A key with a low spending limit is a good idea.

Methods

Methods and decisions

How the numbers on this site were made, what they assume and where they fall short, with the decisions behind them written down. The same documents live in the repository's docs/ folder.

Methods

This page sets out where the data comes from, what was done to it, how the results were evaluated, what the analysis assumes and where it falls short. The 2021 notebook's own cleaning and model are documented rule by rule on the method page. The decision records linked below explain the larger choices.

Data provenance

SourceUsed forVersion and notes
NYC TLC yellow-taxi trip records, 2019every trip-level numberTLC's 2022 Parquet re-issue, 12 monthly files, 84,598,444 rows. The 2021 notebook read the original CSVs, which TLC no longer publishes.
NOAA GHCN-Daily, Central Park stationdaily weather features and the rain analysisthe file the 2021 notebook used, from the 2021 project snapshot
NYPD motor-vehicle collisionsdaily collision countsa 2021 BigQuery export that ends on 23 December 2019
NYC Open Data permitted events (bkfu-528j)daily event counts per boroughthe current version, which is about ten times smaller than the 2021 export
TLC taxi-zone lookup and shapefilezone names, boroughs and map polygonsthe 2021 snapshot

scripts/fetch_data.py downloads the trips and extracts the other files. scripts/pipeline.py applies the 2021 rules and writes a DuckDB work file that is not committed. Five later scripts fit the model, run the rigour, effects and data-quality analyses, and build the database. They write the committed outputs in scripts/out/ and the read-only database web/data/analytics.db, which holds aggregates only and no individual trip. The README lists the order to run them in.

Method

  • Cleaning. The 2021 rules run exactly as the notebook ran them, quirks included (DR-001, DR-002). The data-quality report shows what each rule catches and what got through.
  • The 2021 model. An elastic-net linear regression on 579 features, unchanged. The 2026 refit solves the same objective to convergence from exact sufficient statistics.
  • Inference. Penalised coefficients have no honest standard errors, so the evaluation page reports the unpenalised OLS counterpart on the same design with classical, HC3 and day-clustered standard errors. The day-clustered errors matter because weather, events and collisions are shared by every trip on a day.
  • Prediction intervals. Split-conformal intervals with a Mondrian taxonomy by pickup borough and predicted duration (DR-003).
  • Effects. Rain and permitted events are compared through a composition-adjusted duration index, the mean of log(minutes ÷ median minutes of the same route and hour). Rain uses a two-sample comparison and a day-level regression with month, weekday and holiday effects. Events use a matched comparison of borough-days within the same weekday, four-week window and weather. See the effects page.
  • Ask the data. A language model drafts SQL in the visitor's browser with the visitor's key, and the server validates and runs it read-only (DR-004).

Evaluation design

  • Splits. Folds are a deterministic hash of each trip, because Spark's randomSplit cannot be replayed. The temporal hold-out fits on January to October (folds 1 to 9) and tests on all of November and December. The in-period hold-out tests on January to October fold 0.
  • Baselines. Every model is compared with two baselines. One predicts the January to October mean for every trip. The other looks up the January to October median of the same pickup zone, drop-off zone and hour, falling back to the route and then the pickup zone and hour when a cell has fewer than 20 trips.
  • Uncertainty. Hold-out metrics get a cluster bootstrap that resamples whole days (B = 2,000, seed 20190101), because trips on the same day share weather and traffic. Differences between models are paired, using the same resampled days for both. Interval coverage is a proportion of trips, but coverage also moves together within a day, so it gets the same day bootstrap rather than a Wilson interval, which would be 6 to 12 times too narrow. scripts/rigour.py ports the website's random number generator, so its intervals match the website's bootstrap exactly. Text-to-SQL accuracy, where questions are independent, gets Wilson intervals. The rain regression uses HC3 standard errors with t-based intervals, checked against Newey–West errors with 7 lags because neighbouring days are correlated. The matched event comparison resamples dates, because several boroughs can be event-heavy on the same date. Effect sizes are reported as Hedges' g or d_z alongside the intervals.
  • Text-to-SQL. The 24 evaluation questions and their reference queries were written before any model was run. The domain notes in the "described" prompt were written with these questions in view, and 7 questions depend on a fact a note states, so accuracy is also reported on the other 17. A model passes when its result contains the reference result, with columns matched by value, row order checked only for rankings, and numbers compared to six significant figures. Strict accuracy also requires no extra columns. Two runs are compared with an exact McNemar test and a paired bootstrap interval.
  • Checked code. The TypeScript statistics helpers in web/src/lib/stats/ are unit-tested against values computed with scipy, statsmodels and R (scripts/stats_reference.py). The sparse NumPy code that computes HC3 and clustered standard errors over 75 million rows is checked against statsmodels on a dense test problem before every run.

Assumptions

  • Calibration and test trips are exchangeable relative to the fixed model. This holds for random 2019 trips and fails for later months, which is why coverage is also reported on November and December.
  • Days are independent units for the cluster bootstrap. Consecutive days are correlated through weather systems and seasons, so the day-level intervals are, if anything, a little narrow. For the rain regression, Newey–West errors allow for this and barely change the interval.
  • The 2019 median of each route and hour is a fair yardstick for "usual", and trips that share a route and hour are comparable.
  • Central Park's daily precipitation stands in for rain across the city and across the day.

Limitations

  • One year of data cannot separate seasonal drift from the particular character of November and December traffic.
  • The effects are associations after the stated adjustments. Rain and permits were not randomly assigned.
  • The 2021 model has no notion of distance between zones, which caps its accuracy well below a simple lookup table.
  • The events feature cannot be reproduced exactly, because NYC Open Data revised the 2019 events after 2021.
  • The cleaning quirks are kept on purpose, so some implausible trips remain in every number on the site. Their counts are on the data-quality page.
  • Rebuilding the data needs about 20 GB of memory and 40 GB of disk.

What I'd change

  • Replace the point model with a route-by-hour lookup or a gradient-boosted model, and wrap it in the same conformal intervals. That would narrow the intervals far more than any change to the interval method.
  • Refit under corrected cleaning rules as a sensitivity analysis, so the cost of each quirk is a number rather than a caveat.
  • Use a block bootstrap over weeks to allow for correlation between neighbouring days.
  • Grow the text-to-SQL question set to at least 100 questions, with a second person writing reference answers.

Decision records

Each record gives the context, the decision, the options considered, why, what happened (weak numbers included) and what I would change. Past records are never edited. A changed decision gets a new record that supersedes the old one.

  1. DR-001 · Accepted · 2026-10-06Re-run the 2021 pipeline with DuckDB instead of PySparkTranslate every cleaning rule of the 2021 PySpark notebook into DuckDB SQL, run it in single-file uv scripts, and fit and score the model from exact sufficient statistics instead of a Spark session.
  2. DR-002 · Accepted · 2026-10-06Keep the 2021 cleaning thresholds as they ran, and report what they missRun every 2021 cleaning rule exactly as the notebook's code ran it, including the rules whose code disagrees with their comments, and publish a data-quality report of what each rule catches and what got through instead of quietly fixing them.
  3. DR-003 · Accepted · 2026-10-06Split-conformal prediction intervals, Mondrian by borough and predicted durationThe trip estimator shows a split-conformal prediction interval around the 2021 model's prediction, calibrated separately within each pickup borough and band of predicted duration, at 80%, 90% or 95%, with its empirical coverage on held-out trips stated next to it.
  4. DR-004 · Accepted · 2026-10-06"Ask the data" uses the visitor's own key in the browser and validates SQL on the serverThe optional "Ask the data" feature calls the language model from the visitor's browser with the visitor's own API key, shows the proposed SQL for a human to run, edit or discard, validates and runs that SQL read-only on the server, logs every call in the browser, and ships an evaluation harness that measures how often the model is right.

Model card: 2019 NYC yellow-taxi trip-duration regression

This card describes the regression from the 2021 MAST30034 project as it runs on the revived website. The model itself is unchanged since 2021. The evaluation, intervals and caveats below were added in 2026.

Model details

  • Type: linear regression with an elastic-net penalty, fitted with Spark MLlib LinearRegression(maxIter=10, regParam=0.3, elasticNetParam=0.8) in August 2021.
  • Inputs: 579 features. There are 11 numeric columns for weather, permitted events and collisions, plus one-hot encodings of weekday, pickup hour, rate code, passenger count, pickup zone, vendor, drop-off zone and the store-and-forward flag.
  • Output: predicted trip duration in minutes.
  • Coefficients: fold 1 of the notebook's 10-fold cross-validation (coursework/10-folds-linear-regression.csv). The penalty kept 83 of the 579 coefficients.
  • Owner: Sunchuangyu (Rin) Huang. This was an individual university project.
  • Where it runs: in the visitor's browser on /estimate, from web/src/lib/data/model.json.

Intended use

The model shows how a 2021 coursework regression behaved, term by term. It gives a rough idea of how long a 2019 yellow-cab trip between two taxi zones took at a given hour, with a prediction interval.

It is not meant for real-time arrival estimates, fare or pricing decisions, dispatch, or any decision about drivers or passengers. It knows nothing after 2019 and nothing about green cabs or for-hire vehicles.

Training data

  • Trips: NYC Taxi and Limousine Commission yellow-taxi trip records for 2019. In 2021 these came from TLC's monthly CSVs. The revival uses TLC's 2022 Parquet re-issue of the same months, 84,598,444 raw records.
  • Cleaning: four rounds of rules from the 2021 notebook, kept exactly as they ran (see DR-002 and /data-quality). They leave 74,910,889 trips and 74,941,355 model rows, because the shapefile join duplicates two zone names.
  • Joined data: NOAA GHCN-Daily weather for Central Park, NYPD motor-vehicle collisions (a 2021 BigQuery export that ends on 23 December 2019) and NYC Open Data permitted events. The events dataset has been revised since 2021 and is now about ten times smaller.
  • Personal information: TLC publishes trips without driver or passenger identifiers, with locations coarsened to 263 taxi zones. The website ships aggregates only and no individual trip.

Evaluation

TestResultUncertainty
10-fold cross-validation, 2021 notebookR² 0.3665, RMSE 9.175 minfolds range from R² 0.3655 to 0.3676 and RMSE 9.170 to 9.182
2021 coefficients on the revived foldsR² 0.3681, RMSE 9.179 minfolds range from R² 0.3670 to 0.3689
2021 specification refitted on Jan to Oct, tested on Nov to DecRMSE 9.47 min, MAE 6.79 min, R² 0.35395% CIs 9.18 to 9.73, 6.61 to 6.95, and 0.341 to 0.365 (bootstrap over 61 test days, B = 2,000, seed 20190101)
Route-by-hour median lookup, same hold-outRMSE 6.24 min, R² 0.71995% CI 6.00 to 6.47
90% prediction intervals on 7,497,013 held-out 2019 trips90.0% coverage, 26.4 min mean width95% CI 89.9% to 90.1% (bootstrap over 347 test days); 90% in every borough, 88% to 92% in every predicted-duration decile
90% prediction intervals, refitted on Jan to Oct, tested on Nov to Dec88.8% coverage95% CI 88.4% to 89.2% (bootstrap over 61 test days); 12,936,277 test trips

The temporal hold-out and the intervals are produced by scripts/rigour.py. The hold-out bootstrap runs in web/src/lib/holdout.ts, and the coverage bootstrap runs in scripts/rigour.py with the same generator and seed, checked against the website's bootstrap by a test. The full tables are on /evaluation.

Known failure modes

  • It cannot see distance. The model adds a pickup-zone effect to a drop-off-zone effect, so it has no notion of how far apart two zones are. A plain lookup of the median trip time for the same route and hour beats it by 3.2 minutes of RMSE (95% CI 3.1 to 3.3).
  • Long trips are under-predicted. In the top decile of predictions, a symmetric interval sized for the average trip covers only 64% of trips.
  • Outer boroughs get worse predictions. The mean residual is +8.6 minutes for Bronx pickups and +29.6 minutes for Staten Island pickups, against −0.1 for Manhattan, and the residual SD is 1.6 to 3.6 times Manhattan's.
  • Very short trips can be predicted at under a minute or below zero, because the model is a straight line.
  • The data has gaps. 1 to 20 January 2019 were removed by cleaning, collisions stop on 23 December, and days without a permitted event in the pickup borough were dropped in 2021.
  • Seasons drift. On November and December, after fitting on January to October, the RMSE is about 0.34 minutes worse than on a random January to October fold, and interval coverage slips below its target.

Ethical considerations

Zone coefficients describe traffic, distance and demand. They are not a measure of a neighbourhood, and should not be read as one. The model serves outer-borough riders worst, because most of its training trips start in Manhattan (92% of model rows). Any use that affected people outside Manhattan would need a model that is evaluated separately for them. The website shows aggregates only. The environmental cost of the 2026 analysis was a few minutes of computation on one desktop computer.

Recommendations

Read the prediction together with its interval and with the observed median shown beside it. Where a route-hour median from enough trips exists, it is the better estimate. Treat everything here as a description of 2019.

AI use statement

This statement explains where artificial intelligence is used on the NYC Taxi 2019 website and in its repository, what it is allowed to do, and how its use is recorded. It is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with, or certification under, any of them.

Where AI is used on the website

One feature uses a language model. On /ask, "Ask the data" turns a question in plain English into one SQL query over the site's read-only analytics database. The evaluation harness on /ask/eval uses the same model on 24 fixed questions to measure how often it is right. Both are optional. The rest of the site works without them, and none of its charts, numbers or text is generated by a model when someone visits.

What AI does and never does

The model drafts a SQL query, a short explanation and the assumptions it made. Every output is labelled "AI-generated" with the model's name.

The model never runs anything on its own. In "Ask the data", nothing runs until the visitor reviews the query and chooses to run it, edit it or discard it. In the evaluation harness, the visitor starts a run and the harness then runs each generated query to score it. The model cannot write to the database, never sees individual trips (the database holds aggregates only), and makes no decision about anyone.

Keys and data sent to providers

Visitors bring their own key for Anthropic (Claude Haiku 4.5 by default, or Claude Sonnet 5.5) or OpenAI. The key is kept in the browser's sessionStorage, or in localStorage only if the visitor ticks "remember on this device". It can be forgotten at any time from the AI settings dialog. The browser sends it directly to the chosen provider. It is never sent to this site's server, never logged and never committed to the repository.

Each request sends the provider the visitor's question, a description of the database tables (names, columns, descriptions and a few example values) and the instructions for the model. Nothing else is sent unless the visitor types it into the question. The site's own server receives only the SQL the visitor chooses to run. It checks that the query is a single read-only statement whose query plan is not recursive and not estimated to be expensive. It then runs the query on a read-only connection with a memory limit and a 3-second time limit, and refuses results over 1 MB. Each address can send 20 queries a minute.

Human in the loop and the audit log

Every AI call is written to an audit log in the visitor's browser (IndexedDB). Each entry records the time, the feature, the provider, the model requested and the model that answered, the input (the visitor's question and the prompt settings, never the key), the output, the latency, the token usage reported by the provider and the human decision. The decision is accepted, edited, rejected or pending. Failed calls are logged too, as "no output", with any tokens the provider billed, and evaluation calls are "not applicable". Decisions are appended in order and never overwritten, and the log shows the full history. Visitors can view, export (JSON or CSV) and delete the log at /ai-log. This site keeps no copy.

Measuring the model

The evaluation harness compares the model's query results with hand-written reference answers. It reports execution accuracy with a Wilson 95% interval, also on its own for the 17 questions that the prompt's domain notes were not written for, and it compares two runs (two models, or two prompts) question by question with an exact McNemar test and a paired bootstrap interval. No accuracy figure is published here, because each run costs the visitor money and 24 questions give intervals too wide to support a claim. See DR-004.

AI in building the site

The 2026 revival was written with the help of an AI coding assistant (Claude Code), working under my direction and review. The statistics on the site are produced by scripts and code in the repository and checked by unit tests against scipy, statsmodels and R. None of them is a number a model typed in.

Limitations

A language model can write a query that runs and returns a plausible but wrong answer, especially about units, weekday numbering and which table to use. Read the explanation and assumptions, and check important results against /records. The read-only guard limits what a query can do, but it cannot tell whether a query answers the question that was asked.