Distributions and Dependence7 lessons and their Study toolkit.
Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.
The lessonsin the order to take them.
- 4.1Random SamplingDraw reproducible random samples and know when the sampling rule changes what the sample represents. Open
- 4.2Bernoulli and Binomial OutcomesModel one yes/no outcome and the number of successes in a group. Open
- 4.3The Normal DistributionKnow the Gaussian's shape, the 68–95–99.7 rule, and z-scores. Open
- 4.4Residual Shape and the Q-Q PlotDiagnose the shape of observed data or model residuals without assuming that every ML feature must be normally distributed. Open
- 4.5Long Tails and Log TransformsRecognise long-tailed data and test whether a simple transform helps the model or the interpretation. Open
- 4.6The Multivariate NormalExtend the Gaussian to several features, and read its contours. Open
- 4.7Dependence Over TimeRecognise when nearby observations are related, so an ordinary random split or row bootstrap gives misleading uncertainty. Open
Glossaryevery term the module introduces.
Each with the lesson that introduces it.
- eligible unit
- The object that can be selected once for this task; here, one invoice ID with all of its lines. 4.1
- sampling frame
- The eligible units available to draw from: 25,900 invoice IDs. 4.1
- simple random sample
- Every set of this size has equal chance. 4.1
- without replacement
- A selected unit leaves the available set. 4.1
- inclusion probability
- Chance that a specified eligible unit enters the sample: 200 / 25,900 in the simple design. 4.1
- with replacement
- Return the selected unit before the next draw; 200 such draws gave 199 distinct IDs. 4.1
- stratum
- A predeclared, nonoverlapping group used for selection; here United Kingdom, Germany, France and Other. 4.1
- stratified sample
- Separate random draws within predeclared groups. 4.1
- source-population bias
- Systematic difference between the available frame and the population a claim is meant to describe. 4.1
- weighting (narration)
- Combining strata by each stratum's share of the source frame to estimate an overall figure. 4.1
- churn indicator
- One for a recorded churn label, zero otherwise. 4.2
- parameter p
- The assumed chance of outcome one on one modelled trial; here set to 15.71%. 4.2
- Bernoulli trial
- One modelled outcome that is one with probability p and zero otherwise. 4.2
- expected value
- Probability-weighted average across possible outcomes. 4.2
- variance
- Probability-weighted squared distance from the model mean; p(1 − p) for one Bernoulli trial. 4.2
- binomial count
- Number of ones across a fixed number of Bernoulli trials. 4.2
- independent
- One trial's outcome does not alter another's probability. 4.2
- identically distributed
- Every trial uses the same outcome probabilities. 4.2
- binomial distribution
- Probability pattern of the count across a fixed number of independent Bernoulli trials. 4.2
- binomial mean
- n times the one-trial probability p. 4.2
- standard deviation
- Square root of variance, in the count's units: √(n p (1 − p)). 4.2
- ALT
- Alanine aminotransferase, a laboratory measurement reported in U/L. 4.3
- observed distribution
- Counts or fractions of recorded values. 4.3
- mean, μ
- Arithmetic average used as the fitted curve's centre. 4.3
- standard deviation, σ
- Typical scale of deviations, in the measurement's units. 4.3
- normal distribution
- Symmetric distribution specified by mean and SD. 4.3
- probability density function, PDF
- Curve whose interval areas give probabilities. 4.3
- cumulative distribution function, CDF
- Probability at or below a threshold. 4.3
- tail probability
- Probability beyond a chosen boundary; 2.275% above z = 2. 4.3
- z score
- Distance from the mean measured in SD units. 4.3
- standard normal
- Normal model with mean zero and SD one. 4.3
- normal approximation
- Using a normal model where observed shape is close enough for a stated purpose. 4.3
- predictor
- Recorded input used to calculate a fitted value; here, trip distance. 4.4
- ordinary least squares (narration)
- Chooses the intercept and slope that minimize the summed squared vertical errors. 4.4
- intercept
- Fitted duration at distance zero. 4.4
- slope
- Change in fitted minutes per additional mile. 4.4
- fitted value
- Prediction from the estimated line. 4.4
- residual
- Observed outcome minus fitted value. 4.4
- normal error shape
- A model assumption about conditional errors around fitted values. 4.4
- Q-Q plot
- Paired quantiles from observed data and a reference distribution. 4.4
- normal quantile
- Value at a stated cumulative probability in a normal model. 4.4
- heavy upper tail
- More or larger high values than the chosen reference predicts. 4.4
- positive skew
- A distribution with a longer or heavier right tail. 4.4
- residual SD
- Empirical spread of observed minus fitted durations: about 7.71 minutes here. 4.4
- heteroscedasticity
- Error spread changes with predictor value. 4.4
- constant error spread
- Similar conditional residual variability across predictor values. 4.4
- invoice value
- Sum of included line quantities times their unit prices for one invoice. 4.5
- right skew
- Observations extend farther toward large values than small values. 4.5
- long right tail
- Relatively few observations occupy a wide high-value range. 4.5
- log transform
- Applying a logarithm to each declared target value before fitting. 4.5
- transformed scale
- Units of log(1 + pounds), not pounds. 4.5
- inverse transform
- Operation that maps a transformed value back to the original scale; here
expm1. 4.5 - ordinary least squares
- Coefficients chosen to minimize summed squared deviations on the fitted target scale. 4.5
- held-out error
- Prediction error measured on rows withheld from fitting. 4.5
- MAE
- Mean absolute error, average absolute prediction miss in the original unit. 4.5
- RMSE
- Square root of mean squared prediction error, in the original unit. 4.5
- multivariate normal
- A normal probability model for a vector of jointly varying numeric features. 4.6
- mean vector
- One average coordinate for each feature. 4.6
- sample covariance
- Average centred cross-product, using one less than the sample count. 4.6
- covariance matrix
- All feature variances on the diagonal and pairwise covariances off it. 4.6
- Pearson correlation
- Covariance divided by the product of the two standard deviations. 4.6
- density contour
- Locations with equal fitted probability density. 4.6
- Mahalanobis distance
- Distance from the fitted centre after accounting for covariance. 4.6
- probability ellipse (narration)
- The region inside a squared Mahalanobis radius that holds a stated share of the model's probability. 4.6
- daily series
- One measured value for each calendar day. 4.7
- time dependence
- The value on a date is associated with values on other dates. 4.7
- time series
- Observations indexed in time order. 4.7
- lag
- A fixed number of time steps between paired observations. 4.7
- autocorrelation
- Association between a series and a lagged copy of itself. 4.7
- baseline
- A simple prediction rule used for comparison; here, the training mean for each weekday. 4.7
- absolute error
- Distance between an observed count and its prediction, in the same units. 4.7
- MAE
- Mean of absolute prediction errors, in trips per day. 4.7
- chronological holdout
- A later test period separated from earlier training dates. 4.7
- walk-forward evaluation
- Repeating a time-ordered test as the training cutoff advances. 4.7
- independent observations
- Observations whose joint variation is not tied by the assumed sampling process. 4.7
- temporal leakage
- Use of future information during training or selection for an earlier forecast. 4.7
Module quiz15 questions across it all.
Take it after the last lesson. Your first pick on each question is the one that counts.
Project: A One-Page Retail Sampling-and-Value Briefbuild it without a template.
Stated as a problem, with no step-by-step instructions. Working out the steps is the point.
The task
Published with the module, after 4.5 (the optional 4.4 and 4.6 are not
needed). A retailer wants two things from its invoice records: a quality
review of 200 invoices that covers every country group, and a first check of
whether invoice values can be predicted when an invoice is opened. Your job is
a one-page brief that audits both with the ideas of Module 4, states every
inclusion, weighting and value rule, and says plainly what the held-out errors
support. It is stated as a problem, with no template and no step-by-step
instructions. 05_module_project.ipynb checks the core values; a
worked version is published separately.
The data
The pinned UCI Online Retail archive (online+retail.zip, member Online Retail.xlsx), downloaded and checked by the notebook's first cell: 541,909
invoice lines from 2010-12-01 to 2011-12-09. One row is a line, not an
invoice. Cite the dataset as the learner pack README shows (CC BY
4.0).
Declared before counting
| item | value |
|---|---|
| sampling unit | one invoice ID, cancellations included, with all of its lines |
| strata | recorded country: United Kingdom, Germany, France, and Other for every other country |
| designs | a simple random sample of 200 IDs without replacement, then 50 IDs per stratum without replacement, both from one np.random.default_rng(20260926) generator, in that order |
| audited quantity | the share of invoice IDs that are cancellations (numbers starting with "C") |
| value target | the positive gross invoice value: non-C invoices, positive quantity and price, Quantity × UnitPrice summed by invoice |
| prediction split | invoices opened 2011-01-01 to 2011-09-30 for fitting, 2011-10-01 to 2011-11-30 for scoring |
| predictors | opening country (a United Kingdom flag), weekday and hour, known when an invoice is opened |
- The frame and the two samples. The frame size by stratum, and both samples' counts by stratum. Say which design the review should use, and why.
- Inclusion probabilities and weights for the 50-per-stratum design, one row per stratum.
- The audited share, four ways: the frame's cancellation share, the simple sample's estimate, the stratified sample's unweighted estimate, and its weighted estimate. Say which estimates describe the frame, and why the weighted one still differs from it.
- Two value definitions. Count the signed invoice totals that are zero or negative, and the positive gross invoices, and say which definition the prediction uses and why the other cannot enter a log.
- Raw and log targets on the later period. Fit both on the same predictors and training invoices, score both in pounds on the same later invoices with MAE and RMSE, and give the share of later invoices where the log fit is closer. Declare the error measure your decision uses before you compare.
- What the errors support, in at most four sentences: which fit your measure chooses, what a simple baseline does, and what neither fit establishes.
The brief fits on one page: two small tables and no more than eight sentences. It uses no hypothesis tests or confidence intervals; those come later in the course.
Notebooksthat check your answers.
Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.
- Random sampling and count modelsLessons 4.1 and 4.2 Courses plan
- The normal distribution and residual shapeLessons 4.3 and 4.4 Courses plan
- Long tails, log transforms and the multivariate normalLessons 4.5 and 4.6 Courses plan
- Dependence over timeLessons 4.7 Courses plan
- Stats Module 4 project · A retail sampling-and-value briefThe module project, with checks Courses plan
Referencefor revising and for interviews.
The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.