Statistics · Module 2

EDA, Evidence and Misleading Statistics8 lessons and their Study toolkit.

Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.

8of 8 lessons ready
55quiz questions
75key terms
6self-checking notebooks

The lessonsin the order to take them.

8 lessons ready to study
  1. 2.1EDA Is a Sequence of DecisionsTurn a raw table into a short list of verified observations and the next questions to test. Open
  2. 2.2How Charts Can MisleadDetect when visual choices change the apparent size or shape of an effect without changing the underlying observations. Open
  3. 2.3Denominators, Base Rates and Missing GroupsCheck whether the reported numerator and denominator refer to the population the claim is about. Open
  4. 2.4Simpson's Paradox and Subgroup ComparisonsExplain how an aggregate comparison can point in a different direction from the comparisons within relevant groups. Open
  5. 2.5Selective Reporting and SurvivorshipRecognise a claim selected because it looked strong, and a sample that includes only cases that survived the process of being observed. Open
  6. 2.6What the Evidence Can ClaimSeparate description, prediction and causal effect in a real-world report. Open
  7. 2.7Project Lab: Audit a Business ClaimReproduce, diagnose and rewrite a misleading but realistic statistical claim from a chart and its underlying data. Open
  8. 2.8Project Lab: End-to-End EDA for an ML DecisionProduce a reproducible data report that changes a modelling decision, not a gallery of attractive plots. Open

Glossaryevery term the module introduces.

Each with the lesson that introduces it.

exploratory data analysis (EDA)
Inspecting recorded data, its quality and its patterns before drawing conclusions. 2.1
unit of analysis (narration)
Whatever is counted once for the stated result; here, an included invoice on its recorded day. 2.1
schema
Field names and stored data types in a table, as df.info() shows them. 2.1
missingness
Absent recorded values in a stated field and set of rows: 135,080 lines have no CustomerID. 2.1
duplicate record
A row matching an earlier row on the compared fields: 5,268 exact repeats, to inspect, not delete silently. 2.1
data-quality check
A test for missing, repeated or invalid recorded values before a declared use. It records an issue; the decision about it comes after. 2.1
univariate view (narration)
A view of one variable on its own, such as the distribution of daily invoice counts. 2.1
bivariate view (narration)
A view of two variables together, such as each daily count against its date. 2.1
subgroup view (narration)
The same measure split by a declared category, such as recorded country. 2.1
evidence log (narration)
A record of each observation, a plausible alternative explanation and the next check. 2.1
unit
The quantity a chart's axis measures: pounds of gross positive invoice value, or a count of invoices. 2.2
relative change
Difference divided by the starting value: 22.57% between the weeks starting 2011-10-31 and 2011-11-07. 2.2
axis baseline
The value at the start of a chart axis. 2.2
truncated axis
An axis whose displayed range omits lower values. 2.2
bar encoding
Magnitude represented by length from the baseline, which is why a bar axis starts at zero. 2.2
displayed time window
The dates included on a chart. 2.2
incomplete interval
A displayed time bin extending beyond the source's recorded period, such as the week containing 2011-12-09. 2.2
aggregation window
The span of source dates combined into one reported value. 2.2
rate per week
An amount divided by the number of complete weeks it represents. 2.2
smoothing
Replacing nearby values with a declared local summary, such as a three-week moving average. 2.2
moving average (narration)
The mean of each value and its neighbors; a centered three-week one uses the week and the weeks either side. 2.2
frequency density
Bin count divided by bin width, used when histogram bins have unequal widths. 2.2
dual axes
Separate vertical scales placed on one plot; either can be rescaled until two lines look aligned. 2.2
source population
The recorded cases a summary begins with: 30,000 existing clients. 2.3
base rate
Event fraction in a stated starting population: 6,636 / 30,000 = 22.12%. 2.3
simulation (narration)
An explicit teaching rule or process applied to data and labelled as such; here the screen PAY_0 <= 0, not an actual lender rule. 2.3
selection
A rule determining which source rows enter a reported subset. 2.3
unobserved outcome
A result absent under the decision being evaluated, such as a rejected applicant's default. It is not a zero. 2.3
eligible denominator
The cases allowed into the stated rate under its selection rule: the simulated approvals. 2.3
group size, n
The number of cases in a stated group. 2.3
approval rate
Approved count divided by the stated group considered by the rule. 2.3
reported rate
A numerator and denominator plus the rule defining who is included. 2.3
teacher audit (narration)
Revealing, in a labelled teaching step, outcomes that the simulated process hides and a real lender could never recover. 2.3
population change
A changed set of cases in a summary. 2.3
denominator mismatch
A numerator and denominator that refer to different eligible sets, such as 3,207 approved defaults over all 30,000 rows. 2.3
missing outcome group
Cases whose result cannot be observed under the stated process. 2.3
subgroup
Records selected by a declared category, such as department. 2.4
denominator
All applicants in the stated subgroup. 2.4
Simpson's paradox
An aggregate comparison can reverse when relevant subgroup composition is accounted for. 2.4
composition
How each group's applicants are distributed across departments. 2.4
weighted average
Subgroup rates combined according to declared subgroup shares. 2.4
standardization
Recomputing group rates under one declared subgroup mix. 2.4
causal claim
A claim that a change in one factor produced an outcome change. This table alone does not identify one. 2.4
selective reporting
Showing a result because it looked favorable while hiding the search that selected it. 2.5
winner's curse
A selected maximum tends to overstate the underlying performance of identical candidates. 2.5
follow-up
A later observation of the originally included units. 2.5
intersection
Units meeting both stated conditions, such as clicked earlier and active at follow-up. 2.5
survivorship
Analyzing only units that remain observable after a selection process. 2.5
guardrail metric
A predeclared outcome checked for possible harm while pursuing the primary goal, measured on everyone assigned. 2.5
early stopping (narration)
Ending a test because of the results seen so far, which makes the stopping time part of the search. 2.5
observed outcome
A value recorded or computed from the declared source rows. 2.6
before/after comparison
A measured difference across two time windows. 2.6
causal effect
An outcome difference attributable to an intervention under a valid comparison. 2.6
prediction question (narration)
Whether past values help forecast new cases, evaluated on cases kept out of model fitting. 2.6
random assignment
Using chance to allocate units to comparison groups. 2.6
assignment unit
The entity randomized and followed for the outcome: a customer, not an invoice. 2.6
control
The comparison condition without the proposed intervention. 2.6
treatment
The condition receiving the proposed intervention. 2.6
binary outcome
An outcome recorded as one of two declared values, such as purchase or no purchase in a fixed period. 2.6
treatment effect
An outcome difference attributable to treatment under a valid comparison. 2.6
uncertainty
How much an estimate may vary across comparable samples. 2.6
practical effect size
The magnitude relevant to the decision. 2.6
unit of analysis
The entity counted once for the stated result: one gross positive invoice. 2.7
declared filter
An explicit rule selecting rows for a calculation. 2.7
percentage change
Difference divided by the starting value, expressed per hundred: 21.57% for 255 to 310. 2.7
truncated axis
An axis whose shown range starts above the quantity's zero point, here a constructed floor of 240. 2.7
rate denominator
The group total used beneath the numerator for a stated comparison: that week's distinct active identified customers. 2.7
missing identifier
A recorded event whose entity link is absent: an invoice with no customer ID. 2.7
time window
The exact interval of records included in a comparison. 2.7
target
The recorded outcome a future model is intended to predict: next-month default, coded 0 or 1. 2.8
holdout
Rows reserved from model-directed exploration for later evaluation: the 6,000 pinned IDs. 2.8
data-quality check
A test of recorded fields against the intended analysis. 2.8
base rate
Observed frequency of the target outcome in a stated set of rows: 5,317 / 24,000 = 22.15%. 2.8
feature code
A stored numeric value standing for a documented category or status. 2.8
EDA report
Reproducible observations linked to explicit limits and next checks. 2.8

Module quiz15 questions across it all.

Take it after the last lesson. Your first pick on each question is the one that counts.

Project: A One-Page, Training-Only Credit EDA Reportbuild it without a template.

Stated as a problem, with no step-by-step instructions. Working out the steps is the point.

The task

Published with the module, after 2.8. A modelling team will later evaluate a next-month default model for existing credit card clients. Before anyone trains or scores a model, they need a one-page EDA report that says which findings in the source should change what they check next. Your job is that report. It is stated as a problem, with no template and no step-by-step instructions. 06_module_project.ipynb checks the core values; a worked version is published separately.

The data and the sealed rows

The pinned UCI Default of Credit Card Clients archive (default+of+credit+card+clients.zip, member default of credit card clients.xls, header on the second row), downloaded and checked by the notebook's first cell. One row is one existing client, with a recorded binary outcome default payment next month. Cite the dataset in your report: Yeh, I-C. (2009). Default of Credit Card Clients [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C55S3H (CC BY 4.0).

6,000 client IDs are reserved for the later evaluation. They are the last 6,000 positions of np.random.default_rng(20260926).permutation(sorted_IDs), and the sorted list has the SHA-256 recorded in the module README. The notebook rebuilds the list and checks that hash. Remove those rows before you open any column. Every number, plot and rate in your report comes from the remaining 24,000 training clients; no held-out count, distribution or rate appears anywhere.

Your report must:

  1. State the source, the unit, the target and the split. Name the file, one existing client per row, the recorded next-month default as the target, and the 24,000 / 6,000 split, with the assertion that the two ID sets are disjoint and together cover every source ID.
  2. Disclose prior exposure. Earlier videos, Stats 1.8 and the 2.3 simulation, showed whole-file summaries before this split existed. Say that the held-out rows are sealed for this project's workflow, not unseen across the course.
  3. Give the outcome balance: defaults, nondefaults and the training base rate, all over the same denominator.
  4. Audit the data quality: missing cells, then every recorded code of EDUCATION, MARRIAGE and PAY_0 against the source description. Count each code the description does not define, and do not translate any undocumented code into a status.
  5. Show two numerical fields and one pair view: the distributions of AGE and LIMIT_BAL on their own scales, and one pair plot of the two from a labelled display sample, with Pearson r computed on all training rows.
  6. Compare declared groups with their denominators: the default rate by recorded SEX code, and by the age bands 21–40, 41–60 and 61+, declared before you look at any rate. Print every group's count beside its rate.
  7. Write three next checks, each linked to its evidence: at least one documentation question and one subgroup limitation, each written as observation → question for the modelling team → next verification.
  8. State what the report cannot decide: it chooses no model, estimates no performance, and establishes no causal effect of any recorded field.

Notebooksthat check your answers.

Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.

  • Stats 2.1, 2.2 and 2.7 · EDA and a chart auditLessons 2.1 and 2.2 and 2.7 Courses plan
  • Denominators, base rates and missing groupsLessons 2.3 Courses plan
  • Simpson's paradox and subgroup comparisonsLessons 2.4 Courses plan
  • Selective reporting and survivorshipLessons 2.5 Courses plan
  • What the evidence can claimLessons 2.6 Courses plan
  • Stats Module 2 project · A training-only credit EDA reportThe module project, with checks Courses plan

Referencefor revising and for interviews.

The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.