Estimation5 lessons and their Study toolkit.
Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.
The lessonsin the order to take them.
- 5.1Samples, Sampling Distributions and the Central Limit TheoremSee how an estimate changes across samples, without requiring a formal central-limit-theorem proof. Open
- 5.2Maximum Likelihood EstimationChoose the probability that best explains observed yes/no outcomes, then connect it to logistic-regression log loss. Open
- 5.3Confidence IntervalsReport an estimate with its uncertainty, and say what the interval does and does not mean. Open
- 5.4The BootstrapGet an interval for any statistic by resampling the data you have. Open
- 5.5Bayesian Estimation and Credible IntervalsUpdate uncertainty about an unknown rate as new outcomes arrive, and distinguish a credible interval from a confidence interval. Open
Glossaryevery term the module introduces.
Each with the lesson that introduces it.
- population
- All units in the declared target frame: here, the 19,960 gross positive invoice IDs. 5.1
- sampling unit
- One eligible invoice ID, the unit a draw selects. 5.1
- population mean
- Average over every unit in the declared frame: £534.40. 5.1
- sample
- Selected units from the declared population. 5.1
- sample statistic
- A number calculated from the selected sample, such as its mean. 5.1
- sampling without replacement
- One ID appears at most once in a draw; the full frame is available again for the next draw. 5.1
- sampling distribution
- Distribution of a statistic across repeated samples under one design. 5.1
- standard error
- Spread of a statistic across repeated samples. 5.1
- empirical standard error (narration)
- The SD of the simulated sample means: £370.99 at 25 IDs. 5.1
- finite-population correction
- Reduction in sampling spread when draws use a substantial fraction of a finite frame: √((N − n) / (N − 1)). 5.1
- central limit theorem
- Under suitable conditions, standardized sample means approach a normal distribution as sample size increases; it gives no fixed sample size. 5.1
- Bernoulli outcome
- One recorded zero or one under this model. 5.2
- parameter
- Numerical value chosen within a model, such as the churn probability p. 5.2
- likelihood
- Model probability of the observed labels, treated as a function of the candidate parameter. 5.2
- log likelihood
- Sum of log probabilities assigned to observed outcomes. 5.2
- maximum likelihood estimate
- Parameter value where the observed-data likelihood is largest. 5.2
- p-hat
- Estimate of p from these observed rows: 495 / 3,150. 5.2
- negative log likelihood
- Log likelihood multiplied by minus one. 5.2
- log loss
- Mean negative log likelihood per observed row. 5.2
- nat
- Unit produced when information is measured with natural logarithms. 5.2
- feature
- Recorded input used to vary a model's prediction. 5.2
- linear score (narration)
- An intercept plus a coefficient times the feature, before the sigmoid. 5.2
- sigmoid
- Function mapping any real score into a value between zero and one. 5.2
- logistic regression
- Model that maps a feature-dependent linear score to a probability. 5.2
- intercept
- Score when x is zero. 5.2
- coefficient
- Score change when x rises by one. 5.2
- cancellation record
- An invoice ID with a C prefix; refund status is not established. 5.3
- threshold
- A predeclared value used for this review decision. 5.3
- point estimate
- One statistic computed from one sample. 5.3
- confidence interval
- Bounds produced by a sampling procedure. 5.3
- Wilson confidence interval
- A score-based interval for a binary share. 5.3
- confidence level
- The intended long-run coverage of a procedure under its assumptions. 5.3
- coverage
- Fraction of repeated intervals containing a fixed target. 5.3
- sample mean
- Sum of sampled numeric values divided by sample size. 5.3
- t interval
- A mean interval using sample spread and a t cutoff; its coverage depends on approximation conditions. 5.3
- paired comparison
- Two results measured on the same observational units. 5.4
- ROC-AUC
- Rank-based score comparing positive and negative labels: the share of positive–negative pairs where the positive scores higher, ties counting one half. 5.4
- validation score
- Performance calculated on data withheld from fitting. 5.4
- observed difference
- B's score minus A's score on the original validation set. 5.4
- bootstrap (narration)
- Drawing new samples, with replacement, from the one observed set itself. 5.4
- bootstrap replicate
- One resampled dataset, drawn with replacement from the observed rows. 5.4
- paired resampling
- Applying one draw of unit indices to every measurement on those units. 5.4
- bootstrap distribution
- The collection of a statistic from many bootstrap replicates. 5.4
- percentile bootstrap interval
- Bounds read from specified quantiles of a bootstrap distribution. 5.4
- fixed-model uncertainty
- Variation from resampled validation rows while fitted models stay unchanged. 5.4
- sampling without replacement
- A selected row ID cannot be selected again. 5.5
- Bernoulli rate (narration)
- The probability that a transaction has fraud label one. 5.5
- prior distribution (narration)
- Assigns probability across the possible rates before the reviewed rows supply any evidence. 5.5
- Beta distribution
- A probability distribution over a rate from 0 to 1, with mean a / (a + b). 5.5
- shape parameter
- A number defining an assumed Beta curve before sampled evidence. 5.5
- posterior distribution
- Uncertainty over the rate after updating the declared prior with observed labels. 5.5
- prior sensitivity
- A change in posterior results caused by changing the assumed prior while keeping evidence fixed. 5.5
- credible interval
- A rate range containing a stated share of posterior probability under the stated prior and likelihood. 5.5
- confidence interval
- Interval from a procedure designed for repeated-sample coverage under its assumptions. 5.5
- posterior predictive probability
- The model's chance of one next outcome after averaging over posterior rate uncertainty. 5.5
Module quiz15 questions across it all.
Take it after the last lesson. Your first pick on each question is the one that counts.
Project: A One-Page Retail Uncertainty Briefbuild it without a template.
Stated as a problem, with no step-by-step instructions. Working out the steps is the point.
The task
Published with the module, after 5.3 (5.4 is not needed, and the optional 5.5
is not needed). A retailer asks whether cancellation records for one of its
products are common enough to review, and how precisely a sample can describe
its invoice values. Your job is a one-page brief that answers with the ideas
of Module 5, declares every rule before sampling, and says plainly what the
records cannot establish. It is stated as a problem, with no template and no
step-by-step instructions. 05_module_project.ipynb checks the core
values; a worked version is published separately.
The data
The pinned UCI Online Retail archive (online+retail.zip, member Online Retail.xlsx), downloaded and checked by the notebook's first cell:
541,909 invoice lines from 2010-12-01 to 2011-12-09. One row is a
line, not an invoice. Cite the dataset as the learner pack README
shows (CC BY 4.0).
Declared before sampling
| item | value |
|---|---|
| item | stock code 22423, REGENCY CAKESTAND 3 TIER |
| sampling unit | one invoice ID containing the item, with all of its lines |
| recorded outcome | the invoice number starts with "C": a cancellation record, not a confirmed refund |
| threshold | 3% of the item's invoice IDs |
| design | 200 IDs without replacement, np.random.default_rng(20260926), from the item's IDs sorted by number |
| interval | 95% Wilson interval; the decision is resolved only if the whole interval lies on one side of 3% |
| mean check | 400 IDs without replacement from the gross-positive invoice frame, np.random.default_rng(20260927), with a finite-frame-corrected 95% t interval |
- The item's frame. Its lines, invoice IDs and cancellation-record IDs, counted by ID, not by line.
- One sample, one interval, one decision. The count in the sample, the sample fraction, the Wilson interval and the decision against 3%.
- A check on the procedure. Using the known frame share, repeat the
200-ID draw 3,000 times (
np.random.default_rng(20260927)) and report how many intervals contain it. - A sampled mean against the known frame. The 400-ID sample mean, SD, SE and t interval, beside the frame mean, and one sentence on why 5.3 found this procedure undercovers on this frame.
- Scope. At least three sentences on what the brief does not establish: a refund rate, a customer-level rate, and a future rate.
The brief fits on one page: one small table of counts and intervals, one chart of the interval against the threshold, and no more than eight sentences.
Notebooksthat check your answers.
Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.
- Stats 5.1 and 5.3 · Sampling distributions and confidence intervalsLessons 5.1 and 5.3 Courses plan
- Maximum likelihood estimationLessons 5.2 Courses plan
- The bootstrapLessons 5.4 Courses plan
- Stats 5.5 (optional) · Bayesian estimation and credible intervalsLessons 5.5 Courses plan
- Stats Module 5 project · A retail uncertainty briefThe module project, with checks Courses plan
Referencefor revising and for interviews.
The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.