Guided tour
The project in three short rides
Each video follows one workflow from start to finish, with the step shown on screen and as captions. A Playwright script recorded them from this site and checked every figure it shows on the way (the hourly pickup totals, JFK's median trip at two hours, the prediction and its interval, the rain and event effects with their confidence intervals, the data-quality counts), so a broken feature fails the recording instead of producing a misleading video.
Walkthrough 1 of 3 · /map → /routes
Where and when
The zone choropleth of 74.9 million cleaned 2019 trips: scrub the hour, switch the weekday, recolour by median minutes, open JFK Airport's detail and its 24-hour profile, play the day, then follow JFK's busiest routes.
Steps (transcript)
- 1The zone map: 74.9 million cleaned 2019 trips by taxi zone, here pickups from 6 to 7 pm
- 2Scrub the hour: 4 am is the quietest, 558,493 pickups in 2019 against 4.9 million at 6 pm
- 38 am: the morning peak starts on the Upper East Side
- 4Saturday, 1 am: the East Village and the Lower East Side lead the night
- 5All days at 3 pm, coloured by median minutes: the airports and the outer zones are slowest
- 6Open JFK Airport: its pickups, median trip, 24-hour profile and vendor split
- 7A JFK pickup takes a median 26.2 min at 1 am and 51.8 min at 3 pm
- 8Play the day: the hour advances and the whole map follows
- 9Routes from JFK Airport: the busiest destinations, drawn like subway lines
Walkthrough 2 of 3 · /estimate → /evaluation
Estimate a trip
The 2021 regression running in the browser: pick a pickup and a drop-off zone, a date and an hour, read the prediction with its split-conformal interval next to what riders actually saw, change the interval level and the hour, see every term of the sum, then check how the intervals were validated.
Steps (transcript)
- 1Estimate a trip: the 2021 regression runs in your browser, all 579 features of it
- 2Pick a pickup zone: Penn Station/Madison Sq West
- 3And a drop-off: Times Sq/Theatre District, on Wednesday 9 October 2019 at 5 pm
- 4The 2021 model predicts 14.2 min, with a 90% split-conformal interval of 3.9 to 31.5 min
- 5Riders saw a median of 13.3 min at 5 pm, over 9,856 trips: the dashed blue line
- 6At 95% the interval widens to 3.1 to 36.7 min, and 95.0% of held-out trips like these fell inside
- 7Change the hour to 8 am: 13.2 min, and the interval moves with the prediction
- 8Why that number: a linear model is a sum, and each bar is one part of the trip
- 9How the intervals were checked: coverage on 7.5 million held-out trips, with day-bootstrap CIs
Walkthrough 3 of 3 · /effects → /data-quality
What changes trip time
Does rain or a street event slow a taxi down? A composition-adjusted duration index, the rain regression with bootstrap, HC3 and Newey–West intervals, a matched comparison for permitted events, the caveats, then the data-quality report behind the cleaned data.
Steps (transcript)
- 1What changes trip time? Each trip is compared with its own route and hour
- 2One dot per day: how much longer than usual the day's trips took, against Central Park rain
- 3Rain: the same trips are 3.4% slower on wet days (95% CI 1.9% to 4.8%), over 333 days
- 4Every comparison with its interval: two-sample, regression and rain bins
- 5Permitted events: 31 matched pairs on 16 dates, −0.6% (95% CI −2.8% to +1.8%), no detectable effect
- 6The caveats: observational data, one rain gauge, collisions as a mediator
- 7The data-quality report: what each 2021 cleaning rule removes, and why
- 8Missing values: until 21 January 2019 almost every row lacks the congestion surcharge
- 9Rule by rule: removed in sequence, failing alone and failing only this rule
- 10What got through: vendor 1 leaves the $2.50 surcharge out of the total on 23.4 million trips
Screenshots
Every key feature at a glance
Captured by the same script in light mode at 1440 × 900 (the landing page also in dark mode) and on a 390 px phone. Select one to enlarge it; the arrow keys step through the set and Escape closes it.
Desktop · 1440 × 900
Phone · 390 × 844
How these were made
pnpm showcase runs web/e2e/showcase.spec.ts on the system Chrome. It plays each journey at a human pace with an on-screen caption and a visible cursor, asserts what it shows, and records it at 1280 × 800; ffmpeg then encodes the H.264 videos on this page and the GIFs in the README. The captions, the step lists here and the README walkthrough are the same text. Every figure is either a full count over the 2019 trips or uses the site's fixed bootstrap seed, so a re-run shows the same numbers.
The walkthroughs use no AI. The two “Ask the data” screenshots use no real API key either: the key is a placeholder, every request to the provider is intercepted in the browser, and the reply is a labelled mock whose text starts with “Mocked response for illustration.” and whose model is reported as a mock. The SQL it proposes still goes through this site's real validator and runs on the real read-only database once a person accepts it.