Statistics · Module 1

Describing Data and Reading Statistical Plots10 lessons and their Study toolkit.

Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.

10of 10 lessons ready
65quiz questions
86key terms
8self-checking notebooks

The lessonsin the order to take them.

10 lessons ready to study
  1. 1.1Histograms and DensitySee the shape of one feature, and read a smooth density without mistaking smoothing for observed data. Open
  2. 1.2The CDF, Percentiles and QuantilesRead the fraction of data below any value, and the value below any fraction. Open
  3. 1.3Mean, Median and ModeChoose the right "typical value", and see the mean pulled by one outlier. Open
  4. 1.4Variance, Standard Deviation and IQRMeasure spread, robustly when the data has outliers. Open
  5. 1.5Box Plots and Violin PlotsCompare the distributions of several groups side by side. Open
  6. 1.6Correlation, Confounding and CausationMeasure how two features move together, and know what correlation can never tell you. Open
  7. 1.7Who Is in the Dataset?See how the sampling frame and missing responses can make a precise-looking summary unrepresentative. Open
  8. 1.8Categorical Counts, Rates and Bar PlotsCompare categories without confusing a count, a proportion and a mean. Open
  9. 1.9Scatter, Joint and Density PlotsSee a relationship between two numerical variables, including patterns a single correlation coefficient misses. Open
  10. 1.10Pair Plots and Correlation HeatmapsScreen several relationships while keeping the units, groups and sample size visible. Open

Glossaryevery term the module introduces.

Each with the lesson that introduces it.

bin
A declared interval of values. [5, 10) includes 5 and excludes 10. 1.1
frequency
The number of observations in a bin. 1.1
histogram
Counts observations inside adjacent numeric intervals. 1.1
density
A bar height whose width times height is the bin's fraction. Its unit is per unit of the x axis: per minute for trip durations. 1.1
threshold
The cutoff in the stated decision rule, such as <= 20 minutes. 1.1
kernel density estimate (KDE)
A smoothed density estimated from observations: a small curve (a kernel) is placed at each value and the curves are added. 1.1
rug plot
Marks at observed values, one per recorded duration. 1.1
bandwidth
The amount of smoothing in a KDE. seaborn scales it with bw_adjust. 1.1
rank
An observation's position after sorting, starting at 1 for the smallest. Tied values still get separate ranks. 1.2
CDF
Cumulative probability at or below a value. 1.2
empirical CDF
The observed fraction at or below a value. It steps up by 1 / n at each observation and is flat between them. 1.2
cumulative fraction
The count at or below a value divided by all selected observations. 1.2
quantile
The value at a chosen cumulative fraction. 1.2
percentile
A quantile whose fraction is written as a percentage: the 90th percentile is the 0.9 quantile. (narration) 1.2
quartiles
Quantiles at 25%, 50% and 75%. 1.2
unit of analysis
The thing counted once in a summary. 1.3
gross positive invoice value
The sum of included positive line values for one invoice. It is not net revenue. 1.3
mean
The sum of observed values divided by their count. 1.3
median
The middle ordered value, or the average of the two middle values. 1.3
mode
The most frequent observed value under a stated precision (here, to the penny). 1.3
modal bin
The histogram interval with the highest count under declared edges. 1.3
right skew
A distribution with a longer tail toward larger values. 1.3
outlier
An observation far from most others under a stated comparison. Being unusual does not make it wrong. 1.3
sensitivity check
Recompute after one stated change to see its effect. 1.3
spread
How far observed values extend around a stated center. 1.4
deviation
An observed value minus its group's mean. 1.4
squared deviation
The square of an observed value minus the group mean. 1.4
sample variance
The sum of squared deviations divided by n - 1. Its unit is the original unit squared. 1.4
sample standard deviation
The square root of the sample variance, in the original unit. 1.4
quartile
A value dividing ordered observations at 25%, 50% or 75% under a stated method (NumPy and pandas interpolate linearly by default). 1.4
IQR
The third quartile minus the first quartile: the spread of the middle half. 1.4
sensitivity check
Recompute after one stated change and compare the results. 1.4
box plot
A plot of quartiles, median, whiskers and flagged observations. 1.5
Tukey fence
Q1 - 1.5 × IQR or Q3 + 1.5 × IQR. 1.5
whisker
A line from a box edge to the most extreme observed value within the fence. 1.5
flagged observation
An observed value outside a declared box-plot fence. It is a flag for inspection, not a proven error. 1.5
violin plot
A mirrored estimated density for each group, often with summary marks inside. 1.5
density normalization
The rule mapping estimated density to plotted width. With a separate scale per group, widths do not show group size. 1.5
candidate feature
A recorded input proposed for later model evaluation. 1.6
scatter plot
One point per paired numerical observation. 1.6
linear association
A straight line pattern between two numerical variables. (narration) 1.6
deviation
An observed value minus the declared group's mean. 1.6
covariance
The average signed product of paired deviations. Its unit is the product of the two units. 1.6
Pearson correlation
Standardized linear association between two numerical variables: covariance divided by the two standard deviations. Written r, between -1 and 1. 1.6
subgroup
Records selected by one declared value of another variable. 1.6
observational data
Recorded values without assigning the proposed exposure. 1.6
rate
The selected outcome count divided by the group's total. 1.6
potential confounder
A third variable associated with predictor and outcome that could alter their interpretation. 1.6
causal effect
The outcome change attributable to an intervention under a valid comparison. 1.6
predictive value
Measured improvement on evaluation data that did not guide the exploration. 1.6
target population
The people the intended population claim concerns. 1.7
control total
A published population count used here for comparison. 1.7
sampling frame
The operational set or process from which sample members can be selected. 1.7
nonresponse
A selected person's requested data are not obtained. 1.7
oversampling
Selecting a declared group at a higher rate by design. 1.7
age band
A declared interval of recorded ages. 1.7
survey weight
A released value used with survey guidance to represent the target population in an analysis. 1.7
missingness mechanism
A possible reason a value is missing; the rows alone cannot say which one applies. (narration) 1.7
MCAR (missing completely at random)
Missingness probability independent of observed and unobserved values. 1.7
MAR (missing at random)
Missingness independent of the missing value after conditioning on observed variables. 1.7
MNAR (missing not at random)
Missingness still depends on unavailable values after conditioning on observed information. 1.7
random train/test split
A random partition of the rows already observed. 1.7
category
A named group represented by a recorded value. 1.8
undocumented code
An observed source value without a category label in the declared map. 1.8
frequency
The number of records in a category. 1.8
contingency table
Counts for each combination of two recorded categories. 1.8
numerator
The event count above the division line. 1.8
denominator
All eligible records in the stated group. 1.8
rate
The event count divided by the eligible group count for a stated period. 1.8
grouped bar plot
Side-by-side bars for categories within each group. 1.8
stacked bar plot
Outcome segments placed end to end within a group. Scaled to equal height, it shows fractions and hides group size. 1.8
association
Recorded values vary together in the observed data. 1.8
scatter plot
One point per paired observation on two numerical axes. 1.9
overplotting
Multiple observations cover the same plotted area. 1.9
Pearson correlation
The strength and direction of linear association in paired numerical values. 1.9
hexagonal bin
An equal-area cell that counts paired observations in a two-variable plot. 1.9
density
Observation concentration per unit area under a stated binning or smoothing rule. 1.9
joint distribution
How two variables occur together in paired observations. 1.9
marginal distribution
One variable's distribution after counting across the other variable. 1.9
subgroup
Records selected by a stated field value. 1.9
validation partition
Rows excluded from this lesson's exploratory calculation and later model fitting. 1.10
off-diagonal panel
One pair of different variables in a pair plot. 1.10
pair plot
A grid of pairwise plots with one-variable distributions on the diagonal. 1.10
sparse values
Many recorded zeros with fewer nonzero entries. 1.10
correlation matrix
A table of pairwise correlation coefficients. It is symmetric, with ones on the diagonal. 1.10
heatmap
A matrix whose cell color encodes its numeric value. 1.10

Module quiz15 questions across it all.

Take it after the last lesson. Your first pick on each question is the one that counts.

Project: A One-Page Retail Invoice Reportbuild it without a template.

Stated as a problem, with no step-by-step instructions. Working out the steps is the point.

The task

Published with the module, after 1.9 (1.10 is optional and not needed). A merchant asks what a typical invoice looked like in the UCI Online Retail records, and whether invoices from customers outside the United Kingdom looked different. Your job is a one-page report that answers with the statistics and plots of this module, and says plainly which claims the records cannot support. It is stated as a problem, with no template and no step-by-step instructions. 08_module_project.ipynb checks the core values; a worked version is published separately.

The data

The pinned UCI Online Retail archive (online+retail.zip, member Online Retail.xlsx), downloaded and checked by the notebook's first cell. One raw row is an invoice line, not an invoice and not a customer. The records run from 2010-12-01 to 2011-12-09. Cite the dataset in your report: Chen, D. (2015). Online Retail [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5BW33 (CC BY 4.0).

Your report must:

  1. Declare the unit before any summary. One observation is a gross positive invoice: leave out invoice numbers starting with C and lines with a quantity or unit price of zero or less, then sum Quantity × UnitPrice by InvoiceNo. State the line and invoice counts.
  2. Show the distribution of invoice values in one plot chosen for this shape, with its axes, unit and any bin width or range stated, and say how many invoices lie beyond what the plot shows, if any.
  3. Give a centre and a spread with their counts: the median and the IQR (with Q1 and Q3), beside the mean, and say in one sentence why you lead with the median.
  4. Compare two declared groups: invoices whose recorded country is the United Kingdom against all other countries, with each group's count, median and IQR, in a plot that shows both groups on one value scale.
  5. Inspect the extreme invoice before judging it: find the largest invoice, show its line and the related cancellation record, and report what the pair's quantities add to. Run one sensitivity check and say what it does and does not show.
  6. Inspect what the unit rule let in. The rule excludes numbers starting with C, but every kept invoice number should still be checked. Find the kept invoices whose number is not all digits, show their lines, decide whether each belongs in a report of sales, and state your decision and its effect on the mean and the median, whichever way you decide.
  7. Check the period: count invoices per calendar month and say what the last month's count can and cannot be compared with.
  8. State what the records do not support, in at least four sentences: one each on net revenue, customer spending, causes of a group difference, and current behavior.

Notebooksthat check your answers.

Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.

  • Histograms, density and the ECDFLessons 1.1 and 1.2 Courses plan
  • Invoice centre and spreadLessons 1.3 and 1.4 Courses plan
  • Box plots and violin plotsLessons 1.5 Courses plan
  • Stats 1.6 and 1.10 · Association in the census recordsLessons 1.6 and 1.10 Courses plan
  • Who is in the dataset?Lessons 1.7 Courses plan
  • Categorical counts, rates and bar plotsLessons 1.8 Courses plan
  • Scatter, joint and density plotsLessons 1.9 Courses plan
  • Stats Module 1 project · A retail invoice reportThe module project, with checks Courses plan

Referencefor revising and for interviews.

The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.