Describing Data and Reading Statistical Plots10 lessons and their Study toolkit.
Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.
The lessonsin the order to take them.
- 1.1Histograms and DensitySee the shape of one feature, and read a smooth density without mistaking smoothing for observed data. Open
- 1.2The CDF, Percentiles and QuantilesRead the fraction of data below any value, and the value below any fraction. Open
- 1.3Mean, Median and ModeChoose the right "typical value", and see the mean pulled by one outlier. Open
- 1.4Variance, Standard Deviation and IQRMeasure spread, robustly when the data has outliers. Open
- 1.5Box Plots and Violin PlotsCompare the distributions of several groups side by side. Open
- 1.6Correlation, Confounding and CausationMeasure how two features move together, and know what correlation can never tell you. Open
- 1.7Who Is in the Dataset?See how the sampling frame and missing responses can make a precise-looking summary unrepresentative. Open
- 1.8Categorical Counts, Rates and Bar PlotsCompare categories without confusing a count, a proportion and a mean. Open
- 1.9Scatter, Joint and Density PlotsSee a relationship between two numerical variables, including patterns a single correlation coefficient misses. Open
- 1.10Pair Plots and Correlation HeatmapsScreen several relationships while keeping the units, groups and sample size visible. Open
Glossaryevery term the module introduces.
Each with the lesson that introduces it.
- bin
- A declared interval of values.
[5, 10)includes5and excludes10. 1.1 - frequency
- The number of observations in a bin. 1.1
- histogram
- Counts observations inside adjacent numeric intervals. 1.1
- density
- A bar height whose width times height is the bin's fraction. Its unit is per unit of the x axis: per minute for trip durations. 1.1
- threshold
- The cutoff in the stated decision rule, such as
<= 20minutes. 1.1 - kernel density estimate (KDE)
- A smoothed density estimated from observations: a small curve (a kernel) is placed at each value and the curves are added. 1.1
- rug plot
- Marks at observed values, one per recorded duration. 1.1
- bandwidth
- The amount of smoothing in a KDE. seaborn scales it with
bw_adjust. 1.1 - rank
- An observation's position after sorting, starting at 1 for the smallest. Tied values still get separate ranks. 1.2
- CDF
- Cumulative probability at or below a value. 1.2
- empirical CDF
- The observed fraction at or below a value. It steps up by
1 / nat each observation and is flat between them. 1.2 - cumulative fraction
- The count at or below a value divided by all selected observations. 1.2
- quantile
- The value at a chosen cumulative fraction. 1.2
- percentile
- A quantile whose fraction is written as a percentage: the 90th percentile is the 0.9 quantile. (narration) 1.2
- quartiles
- Quantiles at 25%, 50% and 75%. 1.2
- unit of analysis
- The thing counted once in a summary. 1.3
- gross positive invoice value
- The sum of included positive line values for one invoice. It is not net revenue. 1.3
- mean
- The sum of observed values divided by their count. 1.3
- median
- The middle ordered value, or the average of the two middle values. 1.3
- mode
- The most frequent observed value under a stated precision (here, to the penny). 1.3
- modal bin
- The histogram interval with the highest count under declared edges. 1.3
- right skew
- A distribution with a longer tail toward larger values. 1.3
- outlier
- An observation far from most others under a stated comparison. Being unusual does not make it wrong. 1.3
- sensitivity check
- Recompute after one stated change to see its effect. 1.3
- spread
- How far observed values extend around a stated center. 1.4
- deviation
- An observed value minus its group's mean. 1.4
- squared deviation
- The square of an observed value minus the group mean. 1.4
- sample variance
- The sum of squared deviations divided by
n - 1. Its unit is the original unit squared. 1.4 - sample standard deviation
- The square root of the sample variance, in the original unit. 1.4
- quartile
- A value dividing ordered observations at 25%, 50% or 75% under a stated method (NumPy and pandas interpolate linearly by default). 1.4
- IQR
- The third quartile minus the first quartile: the spread of the middle half. 1.4
- sensitivity check
- Recompute after one stated change and compare the results. 1.4
- box plot
- A plot of quartiles, median, whiskers and flagged observations. 1.5
- Tukey fence
Q1 - 1.5 × IQRorQ3 + 1.5 × IQR. 1.5- whisker
- A line from a box edge to the most extreme observed value within the fence. 1.5
- flagged observation
- An observed value outside a declared box-plot fence. It is a flag for inspection, not a proven error. 1.5
- violin plot
- A mirrored estimated density for each group, often with summary marks inside. 1.5
- density normalization
- The rule mapping estimated density to plotted width. With a separate scale per group, widths do not show group size. 1.5
- candidate feature
- A recorded input proposed for later model evaluation. 1.6
- scatter plot
- One point per paired numerical observation. 1.6
- linear association
- A straight line pattern between two numerical variables. (narration) 1.6
- deviation
- An observed value minus the declared group's mean. 1.6
- covariance
- The average signed product of paired deviations. Its unit is the product of the two units. 1.6
- Pearson correlation
- Standardized linear association between two numerical variables: covariance divided by the two standard deviations. Written
r, between-1and1. 1.6 - subgroup
- Records selected by one declared value of another variable. 1.6
- observational data
- Recorded values without assigning the proposed exposure. 1.6
- rate
- The selected outcome count divided by the group's total. 1.6
- potential confounder
- A third variable associated with predictor and outcome that could alter their interpretation. 1.6
- causal effect
- The outcome change attributable to an intervention under a valid comparison. 1.6
- predictive value
- Measured improvement on evaluation data that did not guide the exploration. 1.6
- target population
- The people the intended population claim concerns. 1.7
- control total
- A published population count used here for comparison. 1.7
- sampling frame
- The operational set or process from which sample members can be selected. 1.7
- nonresponse
- A selected person's requested data are not obtained. 1.7
- oversampling
- Selecting a declared group at a higher rate by design. 1.7
- age band
- A declared interval of recorded ages. 1.7
- survey weight
- A released value used with survey guidance to represent the target population in an analysis. 1.7
- missingness mechanism
- A possible reason a value is missing; the rows alone cannot say which one applies. (narration) 1.7
- MCAR (missing completely at random)
- Missingness probability independent of observed and unobserved values. 1.7
- MAR (missing at random)
- Missingness independent of the missing value after conditioning on observed variables. 1.7
- MNAR (missing not at random)
- Missingness still depends on unavailable values after conditioning on observed information. 1.7
- random train/test split
- A random partition of the rows already observed. 1.7
- category
- A named group represented by a recorded value. 1.8
- undocumented code
- An observed source value without a category label in the declared map. 1.8
- frequency
- The number of records in a category. 1.8
- contingency table
- Counts for each combination of two recorded categories. 1.8
- numerator
- The event count above the division line. 1.8
- denominator
- All eligible records in the stated group. 1.8
- rate
- The event count divided by the eligible group count for a stated period. 1.8
- grouped bar plot
- Side-by-side bars for categories within each group. 1.8
- stacked bar plot
- Outcome segments placed end to end within a group. Scaled to equal height, it shows fractions and hides group size. 1.8
- association
- Recorded values vary together in the observed data. 1.8
- scatter plot
- One point per paired observation on two numerical axes. 1.9
- overplotting
- Multiple observations cover the same plotted area. 1.9
- Pearson correlation
- The strength and direction of linear association in paired numerical values. 1.9
- hexagonal bin
- An equal-area cell that counts paired observations in a two-variable plot. 1.9
- density
- Observation concentration per unit area under a stated binning or smoothing rule. 1.9
- joint distribution
- How two variables occur together in paired observations. 1.9
- marginal distribution
- One variable's distribution after counting across the other variable. 1.9
- subgroup
- Records selected by a stated field value. 1.9
- validation partition
- Rows excluded from this lesson's exploratory calculation and later model fitting. 1.10
- off-diagonal panel
- One pair of different variables in a pair plot. 1.10
- pair plot
- A grid of pairwise plots with one-variable distributions on the diagonal. 1.10
- sparse values
- Many recorded zeros with fewer nonzero entries. 1.10
- correlation matrix
- A table of pairwise correlation coefficients. It is symmetric, with ones on the diagonal. 1.10
- heatmap
- A matrix whose cell color encodes its numeric value. 1.10
Module quiz15 questions across it all.
Take it after the last lesson. Your first pick on each question is the one that counts.
Project: A One-Page Retail Invoice Reportbuild it without a template.
Stated as a problem, with no step-by-step instructions. Working out the steps is the point.
The task
Published with the module, after 1.9 (1.10 is optional and not needed). A
merchant asks what a typical invoice looked like in the UCI Online Retail
records, and whether invoices from customers outside the United Kingdom
looked different. Your job is a one-page report that answers with the
statistics and plots of this module, and says plainly which claims the
records cannot support. It is stated as a problem, with no template and no
step-by-step instructions. 08_module_project.ipynb checks the
core values; a worked version is published separately.
The data
The pinned UCI Online Retail archive (online+retail.zip, member
Online Retail.xlsx), downloaded and checked by the notebook's first cell.
One raw row is an invoice line, not an invoice and not a customer. The
records run from 2010-12-01 to 2011-12-09. Cite the dataset in your report:
Chen, D. (2015). Online Retail [Dataset]. UCI Machine Learning Repository.
https://doi.org/10.24432/C5BW33 (CC BY 4.0).
Your report must:
- Declare the unit before any summary. One observation is a gross
positive invoice: leave out invoice numbers starting with
Cand lines with a quantity or unit price of zero or less, then sumQuantity × UnitPricebyInvoiceNo. State the line and invoice counts. - Show the distribution of invoice values in one plot chosen for this shape, with its axes, unit and any bin width or range stated, and say how many invoices lie beyond what the plot shows, if any.
- Give a centre and a spread with their counts: the median and the IQR (with Q1 and Q3), beside the mean, and say in one sentence why you lead with the median.
- Compare two declared groups: invoices whose recorded country is the United Kingdom against all other countries, with each group's count, median and IQR, in a plot that shows both groups on one value scale.
- Inspect the extreme invoice before judging it: find the largest invoice, show its line and the related cancellation record, and report what the pair's quantities add to. Run one sensitivity check and say what it does and does not show.
- Inspect what the unit rule let in. The rule excludes numbers starting
with
C, but every kept invoice number should still be checked. Find the kept invoices whose number is not all digits, show their lines, decide whether each belongs in a report of sales, and state your decision and its effect on the mean and the median, whichever way you decide. - Check the period: count invoices per calendar month and say what the last month's count can and cannot be compared with.
- State what the records do not support, in at least four sentences: one each on net revenue, customer spending, causes of a group difference, and current behavior.
Notebooksthat check your answers.
Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.
- Histograms, density and the ECDFLessons 1.1 and 1.2 Courses plan
- Invoice centre and spreadLessons 1.3 and 1.4 Courses plan
- Box plots and violin plotsLessons 1.5 Courses plan
- Stats 1.6 and 1.10 · Association in the census recordsLessons 1.6 and 1.10 Courses plan
- Who is in the dataset?Lessons 1.7 Courses plan
- Categorical counts, rates and bar plotsLessons 1.8 Courses plan
- Scatter, joint and density plotsLessons 1.9 Courses plan
- Stats Module 1 project · A retail invoice reportThe module project, with checks Courses plan
Referencefor revising and for interviews.
The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.