Statistics · Module 0

Why Statistics Matters2 lessons and their Study toolkit.

Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.

2of 2 lessons ready
22quiz questions
21key terms
2self-checking notebooks

The lessonsin the order to take them.

2 lessons ready to study
  1. 0.1Why Statistics Matters for ML, AI and Real DecisionsUnderstand what statistics does: describe observed cases and use them carefully to reason about unobserved cases. See why model performance is an estimate about future cases, not a guaranteed score, and why the same reasoning matters outside ML. Open
  2. 0.2From Question to Evidence to ActionFollow the full statistical workflow once before learning its individual tools. Open

Glossaryevery term the module introduces.

Each with the lesson that introduces it.

Statistics
describing observed data and reasoning about unobserved cases. 0.1
Descriptive statistic
a summary of observed data. The median of the January–February trips is one. 0.1
Statistical inference
reasoning from observed cases about unobserved cases, such as using March trips to choose a forecast for trips that have not happened yet. 0.1
Estimate
a value calculated from observed cases. (00.2): the rate measured in the observed period. An estimate can change when the observed cases change. 0.1
Unit
one completed trip. The thing each row, and each count, is about. 0.1
Outcome
trip duration, the recorded quantity the forecast is trying to predict: dropoff time minus pickup time, in minutes. 0.1
Sample
observed recorded trips. 0.1
Target population
future completed trips under similar conditions; the cases a decision is meant to be used on. 0.1
Simple model
a rule, calculated from recorded trips, that gives a prediction for a new trip. Both forecasts in 00.1, one duration for every trip and one duration per pickup zone, are simple models. 0.1
Median
middle value of sorted observations. (00.2): the middle of sorted observations. With an even count, the usual median averages the two middle values. 0.1
Evaluation
measuring a forecast on cases excluded from its calculation. 0.1
Absolute error
size of observed duration minus forecast, sign ignored. 0.1
Mean absolute error
total absolute error divided by evaluated trips; the average size of the error per trip. 0.1
Unit
one completed trip. The thing each row, and each count, is about. 0.2
Population
the trips the decision is meant to cover. In 00.2, completed yellow-taxi trips picked up in Midtown Center on weekdays from 09:00 to 09:59, now and under the same conditions later. 0.2
Scope
the cases and period the result describes. 0.2
Threshold
the duration used to decide whether a trip is on time. In 00.2 it is 20 minutes, tested as duration <= 20. 0.2
Numerator
the trips that meet the condition. 0.2
Denominator
all eligible trips in the stated question. 0.2
Rate
numerator divided by eligible denominator. 0.2
Empirical cumulative distribution
observed fraction at or below each duration. 0.2

Module quiz12 questions across it all.

Take it after the last lesson. Your first pick on each question is the one that counts.

Project: a Dispatcher Decision Briefbuild it without a template.

Stated as a problem, with no step-by-step instructions. Working out the steps is the point.

The task

Question: Did a 20-minute promise meet a 90% service target for trips picked up in Upper East Side South (zone 237) on weekdays from 08:00 to 08:59?

This is 00.2's workflow on a group you have not seen. The zone, days, hour, promise and target are fixed here, before you count anything, so the brief cannot be tuned to its result. Work in 02_module_project.ipynb, which downloads the March 2024 file and checks your values as you go.

Declared before counting

itemvalue
unitone valid, completed yellow-taxi trip
rowsvalid March 2024 trips picked up March 4–31, with the module's filters (duration above 0 and at most 180 minutes; pickup zone 1 to 265)
populationtrips picked up in zone 237, Monday to Friday, pickup hour 08
promiseduration <= 20 minutes
targetat least 90% of covered trips meet the promise
  1. Population and rules. State the table above in your own words, and say what the rows cannot contain (cancelled requests; other periods).
  2. Counts. The number of covered trips, how many met the promise and how many did not. Check that the two add back to the total.
  3. Two charts from the same selected rows. A histogram with five-minute bins from 0 to 80 minutes, and an empirical cumulative distribution, each with a vertical line at 20 minutes and axis labels with units.
  4. The rate and the decision. The on-time rate with its numerator and denominator, compared with the target.
  5. Variation. The rate in each of the four Monday to Sunday weeks, with each week's count.
  6. Comparison. Put your result beside Stats 00.2's Midtown Center result, 2,937 / 3,818 = 76.9%, and say what the comparison does and does not show.
  7. Scope and next check. One paragraph: which trips and period the result describes, what it does not describe, and the next check you recommend.

The brief fits on one page: the two charts, one small table of counts, and no more than six sentences. It uses no hypothesis tests or confidence intervals; those come later in the course.

Notebooksthat check your answers.

Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.

  • From recorded trips to a decisionLessons 0.1 and 0.2 Courses plan
  • Stats Module 0 project · A dispatcher decision briefThe module project, with checks Courses plan

Referencefor revising and for interviews.

The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.