Why Statistics Matters2 lessons and their Study toolkit.
Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.
The lessonsin the order to take them.
- 0.1Why Statistics Matters for ML, AI and Real DecisionsUnderstand what statistics does: describe observed cases and use them carefully to reason about unobserved cases. See why model performance is an estimate about future cases, not a guaranteed score, and why the same reasoning matters outside ML. Open
- 0.2From Question to Evidence to ActionFollow the full statistical workflow once before learning its individual tools. Open
Glossaryevery term the module introduces.
Each with the lesson that introduces it.
- Statistics
- describing observed data and reasoning about unobserved cases. 0.1
- Descriptive statistic
- a summary of observed data. The median of the January–February trips is one. 0.1
- Statistical inference
- reasoning from observed cases about unobserved cases, such as using March trips to choose a forecast for trips that have not happened yet. 0.1
- Estimate
- a value calculated from observed cases. (00.2): the rate measured in the observed period. An estimate can change when the observed cases change. 0.1
- Unit
- one completed trip. The thing each row, and each count, is about. 0.1
- Outcome
- trip duration, the recorded quantity the forecast is trying to predict: dropoff time minus pickup time, in minutes. 0.1
- Sample
- observed recorded trips. 0.1
- Target population
- future completed trips under similar conditions; the cases a decision is meant to be used on. 0.1
- Simple model
- a rule, calculated from recorded trips, that gives a prediction for a new trip. Both forecasts in 00.1, one duration for every trip and one duration per pickup zone, are simple models. 0.1
- Median
- middle value of sorted observations. (00.2): the middle of sorted observations. With an even count, the usual median averages the two middle values. 0.1
- Evaluation
- measuring a forecast on cases excluded from its calculation. 0.1
- Absolute error
- size of observed duration minus forecast, sign ignored. 0.1
- Mean absolute error
- total absolute error divided by evaluated trips; the average size of the error per trip. 0.1
- Unit
- one completed trip. The thing each row, and each count, is about. 0.2
- Population
- the trips the decision is meant to cover. In 00.2, completed yellow-taxi trips picked up in Midtown Center on weekdays from 09:00 to 09:59, now and under the same conditions later. 0.2
- Scope
- the cases and period the result describes. 0.2
- Threshold
- the duration used to decide whether a trip is on time. In 00.2 it is 20 minutes, tested as
duration <= 20. 0.2 - Numerator
- the trips that meet the condition. 0.2
- Denominator
- all eligible trips in the stated question. 0.2
- Rate
- numerator divided by eligible denominator. 0.2
- Empirical cumulative distribution
- observed fraction at or below each duration. 0.2
Module quiz12 questions across it all.
Take it after the last lesson. Your first pick on each question is the one that counts.
Project: a Dispatcher Decision Briefbuild it without a template.
Stated as a problem, with no step-by-step instructions. Working out the steps is the point.
The task
Question: Did a 20-minute promise meet a 90% service target for trips
picked up in Upper East Side South (zone 237) on weekdays from 08:00
to 08:59?
This is 00.2's workflow on a group you have not seen. The zone, days, hour,
promise and target are fixed here, before you count anything, so the brief
cannot be tuned to its result. Work in
02_module_project.ipynb, which
downloads the March 2024 file and checks your values as you go.
Declared before counting
| item | value |
|---|---|
| unit | one valid, completed yellow-taxi trip |
| rows | valid March 2024 trips picked up March 4–31, with the module's filters (duration above 0 and at most 180 minutes; pickup zone 1 to 265) |
| population | trips picked up in zone 237, Monday to Friday, pickup hour 08 |
| promise | duration <= 20 minutes |
| target | at least 90% of covered trips meet the promise |
- Population and rules. State the table above in your own words, and say what the rows cannot contain (cancelled requests; other periods).
- Counts. The number of covered trips, how many met the promise and how many did not. Check that the two add back to the total.
- Two charts from the same selected rows. A histogram with five-minute bins from 0 to 80 minutes, and an empirical cumulative distribution, each with a vertical line at 20 minutes and axis labels with units.
- The rate and the decision. The on-time rate with its numerator and denominator, compared with the target.
- Variation. The rate in each of the four Monday to Sunday weeks, with each week's count.
- Comparison. Put your result beside Stats 00.2's Midtown Center result,
2,937 / 3,818 = 76.9%, and say what the comparison does and does not show. - Scope and next check. One paragraph: which trips and period the result describes, what it does not describe, and the next check you recommend.
The brief fits on one page: the two charts, one small table of counts, and no more than six sentences. It uses no hypothesis tests or confidence intervals; those come later in the course.
Notebooksthat check your answers.
Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.
- From recorded trips to a decisionLessons 0.1 and 0.2 Courses plan
- Stats Module 0 project · A dispatcher decision briefThe module project, with checks Courses plan
Referencefor revising and for interviews.
The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.