Statistics · Module 4

Distributions and Dependence7 lessons and their Study toolkit.

Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.

7of 7 lessons ready
50quiz questions
76key terms
5self-checking notebooks

The lessonsin the order to take them.

7 lessons ready to study
  1. 4.1Random SamplingDraw reproducible random samples and know when the sampling rule changes what the sample represents. Open
  2. 4.2Bernoulli and Binomial OutcomesModel one yes/no outcome and the number of successes in a group. Open
  3. 4.3The Normal DistributionKnow the Gaussian's shape, the 68–95–99.7 rule, and z-scores. Open
  4. 4.4Residual Shape and the Q-Q PlotDiagnose the shape of observed data or model residuals without assuming that every ML feature must be normally distributed. Open
  5. 4.5Long Tails and Log TransformsRecognise long-tailed data and test whether a simple transform helps the model or the interpretation. Open
  6. 4.6The Multivariate NormalExtend the Gaussian to several features, and read its contours. Open
  7. 4.7Dependence Over TimeRecognise when nearby observations are related, so an ordinary random split or row bootstrap gives misleading uncertainty. Open

Glossaryevery term the module introduces.

Each with the lesson that introduces it.

eligible unit
The object that can be selected once for this task; here, one invoice ID with all of its lines. 4.1
sampling frame
The eligible units available to draw from: 25,900 invoice IDs. 4.1
simple random sample
Every set of this size has equal chance. 4.1
without replacement
A selected unit leaves the available set. 4.1
inclusion probability
Chance that a specified eligible unit enters the sample: 200 / 25,900 in the simple design. 4.1
with replacement
Return the selected unit before the next draw; 200 such draws gave 199 distinct IDs. 4.1
stratum
A predeclared, nonoverlapping group used for selection; here United Kingdom, Germany, France and Other. 4.1
stratified sample
Separate random draws within predeclared groups. 4.1
source-population bias
Systematic difference between the available frame and the population a claim is meant to describe. 4.1
weighting (narration)
Combining strata by each stratum's share of the source frame to estimate an overall figure. 4.1
churn indicator
One for a recorded churn label, zero otherwise. 4.2
parameter p
The assumed chance of outcome one on one modelled trial; here set to 15.71%. 4.2
Bernoulli trial
One modelled outcome that is one with probability p and zero otherwise. 4.2
expected value
Probability-weighted average across possible outcomes. 4.2
variance
Probability-weighted squared distance from the model mean; p(1 − p) for one Bernoulli trial. 4.2
binomial count
Number of ones across a fixed number of Bernoulli trials. 4.2
independent
One trial's outcome does not alter another's probability. 4.2
identically distributed
Every trial uses the same outcome probabilities. 4.2
binomial distribution
Probability pattern of the count across a fixed number of independent Bernoulli trials. 4.2
binomial mean
n times the one-trial probability p. 4.2
standard deviation
Square root of variance, in the count's units: √(n p (1 − p)). 4.2
ALT
Alanine aminotransferase, a laboratory measurement reported in U/L. 4.3
observed distribution
Counts or fractions of recorded values. 4.3
mean, μ
Arithmetic average used as the fitted curve's centre. 4.3
standard deviation, σ
Typical scale of deviations, in the measurement's units. 4.3
normal distribution
Symmetric distribution specified by mean and SD. 4.3
probability density function, PDF
Curve whose interval areas give probabilities. 4.3
cumulative distribution function, CDF
Probability at or below a threshold. 4.3
tail probability
Probability beyond a chosen boundary; 2.275% above z = 2. 4.3
z score
Distance from the mean measured in SD units. 4.3
standard normal
Normal model with mean zero and SD one. 4.3
normal approximation
Using a normal model where observed shape is close enough for a stated purpose. 4.3
predictor
Recorded input used to calculate a fitted value; here, trip distance. 4.4
ordinary least squares (narration)
Chooses the intercept and slope that minimize the summed squared vertical errors. 4.4
intercept
Fitted duration at distance zero. 4.4
slope
Change in fitted minutes per additional mile. 4.4
fitted value
Prediction from the estimated line. 4.4
residual
Observed outcome minus fitted value. 4.4
normal error shape
A model assumption about conditional errors around fitted values. 4.4
Q-Q plot
Paired quantiles from observed data and a reference distribution. 4.4
normal quantile
Value at a stated cumulative probability in a normal model. 4.4
heavy upper tail
More or larger high values than the chosen reference predicts. 4.4
positive skew
A distribution with a longer or heavier right tail. 4.4
residual SD
Empirical spread of observed minus fitted durations: about 7.71 minutes here. 4.4
heteroscedasticity
Error spread changes with predictor value. 4.4
constant error spread
Similar conditional residual variability across predictor values. 4.4
invoice value
Sum of included line quantities times their unit prices for one invoice. 4.5
right skew
Observations extend farther toward large values than small values. 4.5
long right tail
Relatively few observations occupy a wide high-value range. 4.5
log transform
Applying a logarithm to each declared target value before fitting. 4.5
transformed scale
Units of log(1 + pounds), not pounds. 4.5
inverse transform
Operation that maps a transformed value back to the original scale; here expm1. 4.5
ordinary least squares
Coefficients chosen to minimize summed squared deviations on the fitted target scale. 4.5
held-out error
Prediction error measured on rows withheld from fitting. 4.5
MAE
Mean absolute error, average absolute prediction miss in the original unit. 4.5
RMSE
Square root of mean squared prediction error, in the original unit. 4.5
multivariate normal
A normal probability model for a vector of jointly varying numeric features. 4.6
mean vector
One average coordinate for each feature. 4.6
sample covariance
Average centred cross-product, using one less than the sample count. 4.6
covariance matrix
All feature variances on the diagonal and pairwise covariances off it. 4.6
Pearson correlation
Covariance divided by the product of the two standard deviations. 4.6
density contour
Locations with equal fitted probability density. 4.6
Mahalanobis distance
Distance from the fitted centre after accounting for covariance. 4.6
probability ellipse (narration)
The region inside a squared Mahalanobis radius that holds a stated share of the model's probability. 4.6
daily series
One measured value for each calendar day. 4.7
time dependence
The value on a date is associated with values on other dates. 4.7
time series
Observations indexed in time order. 4.7
lag
A fixed number of time steps between paired observations. 4.7
autocorrelation
Association between a series and a lagged copy of itself. 4.7
baseline
A simple prediction rule used for comparison; here, the training mean for each weekday. 4.7
absolute error
Distance between an observed count and its prediction, in the same units. 4.7
MAE
Mean of absolute prediction errors, in trips per day. 4.7
chronological holdout
A later test period separated from earlier training dates. 4.7
walk-forward evaluation
Repeating a time-ordered test as the training cutoff advances. 4.7
independent observations
Observations whose joint variation is not tied by the assumed sampling process. 4.7
temporal leakage
Use of future information during training or selection for an earlier forecast. 4.7

Module quiz15 questions across it all.

Take it after the last lesson. Your first pick on each question is the one that counts.

Project: A One-Page Retail Sampling-and-Value Briefbuild it without a template.

Stated as a problem, with no step-by-step instructions. Working out the steps is the point.

The task

Published with the module, after 4.5 (the optional 4.4 and 4.6 are not needed). A retailer wants two things from its invoice records: a quality review of 200 invoices that covers every country group, and a first check of whether invoice values can be predicted when an invoice is opened. Your job is a one-page brief that audits both with the ideas of Module 4, states every inclusion, weighting and value rule, and says plainly what the held-out errors support. It is stated as a problem, with no template and no step-by-step instructions. 05_module_project.ipynb checks the core values; a worked version is published separately.

The data

The pinned UCI Online Retail archive (online+retail.zip, member Online Retail.xlsx), downloaded and checked by the notebook's first cell: 541,909 invoice lines from 2010-12-01 to 2011-12-09. One row is a line, not an invoice. Cite the dataset as the learner pack README shows (CC BY 4.0).

Declared before counting

itemvalue
sampling unitone invoice ID, cancellations included, with all of its lines
stratarecorded country: United Kingdom, Germany, France, and Other for every other country
designsa simple random sample of 200 IDs without replacement, then 50 IDs per stratum without replacement, both from one np.random.default_rng(20260926) generator, in that order
audited quantitythe share of invoice IDs that are cancellations (numbers starting with "C")
value targetthe positive gross invoice value: non-C invoices, positive quantity and price, Quantity × UnitPrice summed by invoice
prediction splitinvoices opened 2011-01-01 to 2011-09-30 for fitting, 2011-10-01 to 2011-11-30 for scoring
predictorsopening country (a United Kingdom flag), weekday and hour, known when an invoice is opened
  1. The frame and the two samples. The frame size by stratum, and both samples' counts by stratum. Say which design the review should use, and why.
  2. Inclusion probabilities and weights for the 50-per-stratum design, one row per stratum.
  3. The audited share, four ways: the frame's cancellation share, the simple sample's estimate, the stratified sample's unweighted estimate, and its weighted estimate. Say which estimates describe the frame, and why the weighted one still differs from it.
  4. Two value definitions. Count the signed invoice totals that are zero or negative, and the positive gross invoices, and say which definition the prediction uses and why the other cannot enter a log.
  5. Raw and log targets on the later period. Fit both on the same predictors and training invoices, score both in pounds on the same later invoices with MAE and RMSE, and give the share of later invoices where the log fit is closer. Declare the error measure your decision uses before you compare.
  6. What the errors support, in at most four sentences: which fit your measure chooses, what a simple baseline does, and what neither fit establishes.

The brief fits on one page: two small tables and no more than eight sentences. It uses no hypothesis tests or confidence intervals; those come later in the course.

Notebooksthat check your answers.

Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.

  • Random sampling and count modelsLessons 4.1 and 4.2 Courses plan
  • The normal distribution and residual shapeLessons 4.3 and 4.4 Courses plan
  • Long tails, log transforms and the multivariate normalLessons 4.5 and 4.6 Courses plan
  • Dependence over timeLessons 4.7 Courses plan
  • Stats Module 4 project · A retail sampling-and-value briefThe module project, with checks Courses plan

Referencefor revising and for interviews.

The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.