Stats 0.1Concept5 parts

Why Statistics Matters for ML, AI and Real Decisions

Understand what statistics does: describe observed cases and use them carefully to reason about unobserved cases. See why model performance is an estimate about future cases, not a guaranteed score, and why the same reasoning matters outside ML.

In this lesson5 parts

  1. 01Statistics, forecasts and models
  2. 02Sample, target population and the median
  3. 03Absolute error and mean absolute error
  4. 04Estimates and statistical inference
  5. 05Reproducible evaluation in pandas

Key terms

The words this lesson introduces, each in one line. The module’s glossary collects them all.

Statistics
describing observed data and reasoning about unobserved cases.
Descriptive statistic
a summary of observed data. The median of the January–February trips is one.
Statistical inference
reasoning from observed cases about unobserved cases, such as using March trips to choose a forecast for trips that have not happened yet.
Estimate
a value calculated from observed cases. (00.2): the rate measured in the observed period. An estimate can change when the observed cases change.
Unit
one completed trip. The thing each row, and each count, is about.
Outcome
trip duration, the recorded quantity the forecast is trying to predict: dropoff time minus pickup time, in minutes.
Sample
observed recorded trips.
Target population
future completed trips under similar conditions; the cases a decision is meant to be used on.
Simple model
a rule, calculated from recorded trips, that gives a prediction for a new trip. Both forecasts in 00.1, one duration for every trip and one duration per pickup zone, are simple models.
Median
middle value of sorted observations. (00.2): the middle of sorted observations. With an even count, the usual median averages the two middle values.
Evaluation
measuring a forecast on cases excluded from its calculation.
Absolute error
size of observed duration minus forecast, sign ignored.
Mean absolute error
total absolute error divided by evaluated trips; the average size of the error per trip.

Quiz 5 questions

Your first pick on each question is the one that counts, and a right one earns a coin. Getting one wrong here is how the lesson sticks.

Practice

Problems to solve in your own notebook. Each states the problem, not the steps: working out the steps is the exercise. Level A applies the lesson, B combines it with earlier ones, C stretches it.

The self-checking notebook for this lesson is From recorded trips to a decision.

Common mistakes

What you will see when it goes wrong, why it happens, and the fix.

Where it’s used

Where this lesson’s ideas turn up in real work.