EDA, Evidence and Misleading Statistics8 lessons and their Study toolkit.
Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.
The lessonsin the order to take them.
- 2.1EDA Is a Sequence of DecisionsTurn a raw table into a short list of verified observations and the next questions to test. Open
- 2.2How Charts Can MisleadDetect when visual choices change the apparent size or shape of an effect without changing the underlying observations. Open
- 2.3Denominators, Base Rates and Missing GroupsCheck whether the reported numerator and denominator refer to the population the claim is about. Open
- 2.4Simpson's Paradox and Subgroup ComparisonsExplain how an aggregate comparison can point in a different direction from the comparisons within relevant groups. Open
- 2.5Selective Reporting and SurvivorshipRecognise a claim selected because it looked strong, and a sample that includes only cases that survived the process of being observed. Open
- 2.6What the Evidence Can ClaimSeparate description, prediction and causal effect in a real-world report. Open
- 2.7Project Lab: Audit a Business ClaimReproduce, diagnose and rewrite a misleading but realistic statistical claim from a chart and its underlying data. Open
- 2.8Project Lab: End-to-End EDA for an ML DecisionProduce a reproducible data report that changes a modelling decision, not a gallery of attractive plots. Open
Glossaryevery term the module introduces.
Each with the lesson that introduces it.
- exploratory data analysis (EDA)
- Inspecting recorded data, its quality and its patterns before drawing conclusions. 2.1
- unit of analysis (narration)
- Whatever is counted once for the stated result; here, an included invoice on its recorded day. 2.1
- schema
- Field names and stored data types in a table, as
df.info()shows them. 2.1 - missingness
- Absent recorded values in a stated field and set of rows: 135,080 lines have no
CustomerID. 2.1 - duplicate record
- A row matching an earlier row on the compared fields: 5,268 exact repeats, to inspect, not delete silently. 2.1
- data-quality check
- A test for missing, repeated or invalid recorded values before a declared use. It records an issue; the decision about it comes after. 2.1
- univariate view (narration)
- A view of one variable on its own, such as the distribution of daily invoice counts. 2.1
- bivariate view (narration)
- A view of two variables together, such as each daily count against its date. 2.1
- subgroup view (narration)
- The same measure split by a declared category, such as recorded country. 2.1
- evidence log (narration)
- A record of each observation, a plausible alternative explanation and the next check. 2.1
- unit
- The quantity a chart's axis measures: pounds of gross positive invoice value, or a count of invoices. 2.2
- relative change
- Difference divided by the starting value: 22.57% between the weeks starting 2011-10-31 and 2011-11-07. 2.2
- axis baseline
- The value at the start of a chart axis. 2.2
- truncated axis
- An axis whose displayed range omits lower values. 2.2
- bar encoding
- Magnitude represented by length from the baseline, which is why a bar axis starts at zero. 2.2
- displayed time window
- The dates included on a chart. 2.2
- incomplete interval
- A displayed time bin extending beyond the source's recorded period, such as the week containing 2011-12-09. 2.2
- aggregation window
- The span of source dates combined into one reported value. 2.2
- rate per week
- An amount divided by the number of complete weeks it represents. 2.2
- smoothing
- Replacing nearby values with a declared local summary, such as a three-week moving average. 2.2
- moving average (narration)
- The mean of each value and its neighbors; a centered three-week one uses the week and the weeks either side. 2.2
- frequency density
- Bin count divided by bin width, used when histogram bins have unequal widths. 2.2
- dual axes
- Separate vertical scales placed on one plot; either can be rescaled until two lines look aligned. 2.2
- source population
- The recorded cases a summary begins with: 30,000 existing clients. 2.3
- base rate
- Event fraction in a stated starting population: 6,636 / 30,000 = 22.12%. 2.3
- simulation (narration)
- An explicit teaching rule or process applied to data and labelled as such; here the screen
PAY_0 <= 0, not an actual lender rule. 2.3 - selection
- A rule determining which source rows enter a reported subset. 2.3
- unobserved outcome
- A result absent under the decision being evaluated, such as a rejected applicant's default. It is not a zero. 2.3
- eligible denominator
- The cases allowed into the stated rate under its selection rule: the simulated approvals. 2.3
- group size, n
- The number of cases in a stated group. 2.3
- approval rate
- Approved count divided by the stated group considered by the rule. 2.3
- reported rate
- A numerator and denominator plus the rule defining who is included. 2.3
- teacher audit (narration)
- Revealing, in a labelled teaching step, outcomes that the simulated process hides and a real lender could never recover. 2.3
- population change
- A changed set of cases in a summary. 2.3
- denominator mismatch
- A numerator and denominator that refer to different eligible sets, such as 3,207 approved defaults over all 30,000 rows. 2.3
- missing outcome group
- Cases whose result cannot be observed under the stated process. 2.3
- subgroup
- Records selected by a declared category, such as department. 2.4
- denominator
- All applicants in the stated subgroup. 2.4
- Simpson's paradox
- An aggregate comparison can reverse when relevant subgroup composition is accounted for. 2.4
- composition
- How each group's applicants are distributed across departments. 2.4
- weighted average
- Subgroup rates combined according to declared subgroup shares. 2.4
- standardization
- Recomputing group rates under one declared subgroup mix. 2.4
- causal claim
- A claim that a change in one factor produced an outcome change. This table alone does not identify one. 2.4
- selective reporting
- Showing a result because it looked favorable while hiding the search that selected it. 2.5
- winner's curse
- A selected maximum tends to overstate the underlying performance of identical candidates. 2.5
- follow-up
- A later observation of the originally included units. 2.5
- intersection
- Units meeting both stated conditions, such as clicked earlier and active at follow-up. 2.5
- survivorship
- Analyzing only units that remain observable after a selection process. 2.5
- guardrail metric
- A predeclared outcome checked for possible harm while pursuing the primary goal, measured on everyone assigned. 2.5
- early stopping (narration)
- Ending a test because of the results seen so far, which makes the stopping time part of the search. 2.5
- observed outcome
- A value recorded or computed from the declared source rows. 2.6
- before/after comparison
- A measured difference across two time windows. 2.6
- causal effect
- An outcome difference attributable to an intervention under a valid comparison. 2.6
- prediction question (narration)
- Whether past values help forecast new cases, evaluated on cases kept out of model fitting. 2.6
- random assignment
- Using chance to allocate units to comparison groups. 2.6
- assignment unit
- The entity randomized and followed for the outcome: a customer, not an invoice. 2.6
- control
- The comparison condition without the proposed intervention. 2.6
- treatment
- The condition receiving the proposed intervention. 2.6
- binary outcome
- An outcome recorded as one of two declared values, such as purchase or no purchase in a fixed period. 2.6
- treatment effect
- An outcome difference attributable to treatment under a valid comparison. 2.6
- uncertainty
- How much an estimate may vary across comparable samples. 2.6
- practical effect size
- The magnitude relevant to the decision. 2.6
- unit of analysis
- The entity counted once for the stated result: one gross positive invoice. 2.7
- declared filter
- An explicit rule selecting rows for a calculation. 2.7
- percentage change
- Difference divided by the starting value, expressed per hundred: 21.57% for 255 to 310. 2.7
- truncated axis
- An axis whose shown range starts above the quantity's zero point, here a constructed floor of 240. 2.7
- rate denominator
- The group total used beneath the numerator for a stated comparison: that week's distinct active identified customers. 2.7
- missing identifier
- A recorded event whose entity link is absent: an invoice with no customer ID. 2.7
- time window
- The exact interval of records included in a comparison. 2.7
- target
- The recorded outcome a future model is intended to predict: next-month default, coded 0 or 1. 2.8
- holdout
- Rows reserved from model-directed exploration for later evaluation: the 6,000 pinned IDs. 2.8
- data-quality check
- A test of recorded fields against the intended analysis. 2.8
- base rate
- Observed frequency of the target outcome in a stated set of rows: 5,317 / 24,000 = 22.15%. 2.8
- feature code
- A stored numeric value standing for a documented category or status. 2.8
- EDA report
- Reproducible observations linked to explicit limits and next checks. 2.8
Module quiz15 questions across it all.
Take it after the last lesson. Your first pick on each question is the one that counts.
Project: A One-Page, Training-Only Credit EDA Reportbuild it without a template.
Stated as a problem, with no step-by-step instructions. Working out the steps is the point.
The task
Published with the module, after 2.8. A modelling team will later evaluate a
next-month default model for existing credit card clients. Before anyone
trains or scores a model, they need a one-page EDA report that says which
findings in the source should change what they check next. Your job is that
report. It is stated as a problem, with no template and no step-by-step
instructions. 06_module_project.ipynb checks the core values; a
worked version is published separately.
The data and the sealed rows
The pinned UCI Default of Credit Card Clients archive
(default+of+credit+card+clients.zip, member default of credit card clients.xls, header on the second row), downloaded and checked by the
notebook's first cell. One row is one existing client, with a recorded binary
outcome default payment next month. Cite the dataset in your report: Yeh,
I-C. (2009). Default of Credit Card Clients [Dataset]. UCI Machine Learning
Repository. https://doi.org/10.24432/C55S3H (CC BY 4.0).
6,000 client IDs are reserved for the later evaluation. They are the last
6,000 positions of np.random.default_rng(20260926).permutation(sorted_IDs),
and the sorted list has the SHA-256 recorded in the module
README. The notebook rebuilds the list and checks that hash.
Remove those rows before you open any column. Every number, plot and rate
in your report comes from the remaining 24,000 training clients; no held-out
count, distribution or rate appears anywhere.
Your report must:
- State the source, the unit, the target and the split. Name the file, one existing client per row, the recorded next-month default as the target, and the 24,000 / 6,000 split, with the assertion that the two ID sets are disjoint and together cover every source ID.
- Disclose prior exposure. Earlier videos, Stats 1.8 and the 2.3 simulation, showed whole-file summaries before this split existed. Say that the held-out rows are sealed for this project's workflow, not unseen across the course.
- Give the outcome balance: defaults, nondefaults and the training base rate, all over the same denominator.
- Audit the data quality: missing cells, then every recorded code of
EDUCATION,MARRIAGEandPAY_0against the source description. Count each code the description does not define, and do not translate any undocumented code into a status. - Show two numerical fields and one pair view: the distributions of
AGEandLIMIT_BALon their own scales, and one pair plot of the two from a labelled display sample, with Pearson r computed on all training rows. - Compare declared groups with their denominators: the default rate by
recorded
SEXcode, and by the age bands21–40,41–60and61+, declared before you look at any rate. Print every group's count beside its rate. - Write three next checks, each linked to its evidence: at least one documentation question and one subgroup limitation, each written as observation → question for the modelling team → next verification.
- State what the report cannot decide: it chooses no model, estimates no performance, and establishes no causal effect of any recorded field.
Notebooksthat check your answers.
Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.
- Stats 2.1, 2.2 and 2.7 · EDA and a chart auditLessons 2.1 and 2.2 and 2.7 Courses plan
- Denominators, base rates and missing groupsLessons 2.3 Courses plan
- Simpson's paradox and subgroup comparisonsLessons 2.4 Courses plan
- Selective reporting and survivorshipLessons 2.5 Courses plan
- What the evidence can claimLessons 2.6 Courses plan
- Stats Module 2 project · A training-only credit EDA reportThe module project, with checks Courses plan
Referencefor revising and for interviews.
The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.