Statistics · Module 3

Probability5 lessons and their Study toolkit.

Watch each lesson, answer its quiz, then do its practice. After the last lesson: the module quiz, the project, and the interview questions.

5of 5 lessons ready
40quiz questions
43key terms
5self-checking notebooks

The lessonsin the order to take them.

5 lessons ready to study
  1. 3.1Probability BasicsReason about chance with outcomes, events and the rules that combine them. Open
  2. 3.2Conditional Probability and IndependenceUpdate a probability given evidence, and tell independent events from mutually exclusive ones. Open
  3. 3.3Bayes' TheoremTurn "how likely is the evidence given the class" into "how likely is the class given the evidence". Open
  4. 3.4Random Variables, Expectation and VarianceDescribe an uncertain quantity by its distribution, its average and its spread. Open
  5. 3.5Monte Carlo SimulationApproximate an expected outcome when several uncertain inputs interact and the calculation is hard to do by hand. Open

Glossaryevery term the module introduces.

Each with the lesson that introduces it.

outcome
One possible result of the selection; here, one recorded transaction. 3.1
sample space
The outcomes under consideration; here, all 284,807 recorded transactions. 3.1
event
A specified set of outcomes, such as F, the 492 rows with Class = 1. 3.1
empirical probability
Event count divided by the recorded total: 492 / 284,807 = 0.173%. 3.1
event A
The rows satisfying a declared condition; here Amount > 100, 56,508 rows. 3.1
complement
Every outcome in the sample space outside an event. 3.1
intersection, or joint event
Outcomes satisfying both events: 130 rows above 100 and labelled fraud. 3.1
joint probability
Fraction of the sample space satisfying both events. 3.1
union
Outcomes satisfying at least one of the events: 56,870 rows. 3.1
with replacement
Each source label remains available for every draw. 3.1
law of large numbers
With repeated independent draws from the same distribution, a running fraction tends toward that distribution's probability. 3.1
conditional probability (narration)
A fraction calculated inside a restricted group: P(A given B) counts the outcomes in both and divides by all of B. P(spam given flagged) = 9 / 108. 3.2
given (narration)
The word that names the group used as the denominator. 3.2
independence
Knowing one event does not change the probability of the other. 3.2
mutually exclusive
Two events with no shared outcomes, such as spam and real. 3.2
feature (narration)
A recorded property of a message, such as whether a certain word appears in it. 3.2
prior
Probability before the new evidence: 1% spam in the constructed inbox. 3.3
likelihood
Probability of this evidence given a specified class: P(flagged given spam) = 90%. 3.3
false-positive rate
Probability of a positive flag when the class is real: 10% in the construction. 3.3
law of total probability
The probability of the evidence is the sum over all disjoint paths that produce it: 108 flags per 1,000 messages, from spam and from real mail. 3.3
posterior
Probability after incorporating the stated evidence: 8.33% after one flag. 3.3
Bayes' theorem (narration)
The update from a prior and the likelihoods in each class to a posterior: likelihood times prior, divided by the total probability of the evidence. 3.3
conditional independence
Within each class, one flag result does not alter another flag's probability. An assumption the construction does not establish. 3.3
review rule
A declared test that assigns each row to review or no review, such as PAY_0 > 0. 3.4
outcome cell
One combination of assigned action and recorded label, such as a missed default. 3.4
cost matrix
An assumed cost for each action × outcome combination: $1,000 per missed default, $100 per reviewed nondefault, $0 otherwise. 3.4
random variable
A numeric value assigned to each possible outcome; here, a randomly chosen client's assigned cost. 3.4
discrete outcome
One of a countable set of values: $0, $100 or $1,000. 3.4
continuous random variable
Under a model, any value in an interval is possible, such as a review time from 0 to 10 minutes. 3.4
uniform distribution
Equal-length intervals have equal probability: 2 to 4 minutes has 20% under the 0–10 model. 3.4
expected value
The probability-weighted average of a random variable's possible values. 3.4
long-run average (narration)
The value a running average of repeated random draws tends toward: the expected value, by the law of large numbers. 3.4
expected cost
Average assigned cost under the stated outcome frequencies and cost matrix: $117.75 for PAY_0 > 0, $97.98 for PAY_0 >= 0. 3.4
empirical variance
The mean of squared distances from the empirical mean, in squared units. 3.4
standard deviation
Square root of variance, in the original units: about $306.02 for the first rule's costs. 3.4
input distribution
The values a simulation is allowed to draw under its model: the 334 recorded daily totals of item 85123A. 3.5
zero recorded sales
No included positive item units in this file on that date (57 days); not zero demand. 3.5
resampling with replacement
Each draw uses the same source values, so repeats are possible. 3.5
Monte Carlo simulation
Repeated random draws and outcome calculations under stated assumptions. 3.5
Monte Carlo error
Difference caused by using finitely many random draws under a fixed model: 30.44% simulated against 29.94% exact. 3.5
stockout
The simulated total exceeds the candidate stock. 3.5
leftover
The candidate stock exceeds the simulated total. 3.5
assumption uncertainty
Uncertainty about whether model inputs and rules describe the real decision; more runs do not reduce it. 3.5

Module quiz15 questions across it all.

Take it after the last lesson. Your first pick on each question is the one that counts.

Project: A One-Page Fraud Review-Queue Briefbuild it without a template.

Stated as a problem, with no step-by-step instructions. Working out the steps is the point.

The task

Published with the module, after 3.4 (3.5 is optional and not needed). A fraud team is choosing a simple amount rule for its manual review queue, and asks what two candidate thresholds would have done on the recorded transactions. Your job is a one-page brief that answers with the probability of Module 3, labels every assumption, and says plainly why the historical fractions cannot promise future performance. It is stated as a problem, with no template and no step-by-step instructions. 05_module_project.ipynb checks the core values; a worked version is published separately.

The data

The pinned ULB credit-card transactions (creditcard.csv, from TensorFlow's public mirror), downloaded and checked by the notebook's first cell: one row per recorded transaction, 284,807 rows, with a Class label (1 for fraud). Amount has no currency in the source description, and Time counts seconds from the first recorded transaction, across about 48 hours. Cite the dataset as the learner pack README shows; its Kaggle licence, DbCL-1.0, is still to be confirmed for any republished excerpt.

Declared before counting

itemvalue
unitone recorded transaction
event FClass == 1, the recorded fraud label
threshold rulesAmount > 50 and Amount > 250, both strict, applied to the same rows
queuethe rows a rule selects, chosen before any label is known
assumed cost1 cost point per reviewed transaction and 100 per missed labelled fraud; points, not money

Your brief must:

  1. State the source, the unit, the event and the two rules, with the assumed cost points labelled as assumptions.
  2. Count each rule's queue: its size and share of the file, the labelled frauds it captures and the ones it misses. Check that the captured and missed counts add up to the file's labelled frauds.
  3. Give two conditional fractions for each rule: P(queued | fraud), the share of labelled frauds the queue captures, and P(fraud | queued), the share of the queue that is labelled fraud. Name each denominator.
  4. Cost both rules in assumed points per recorded transaction, and find the missed-fraud cost at which the two rules would tie. Say what the ranking depends on.
  5. Check whether the history holds still: split the rows at half the recorded Time span and compare each rule's capture rate in the two halves.
  6. Recommend, with limits: which rule, if either, you would pilot, why neither catches most frauds, what a review capacity would change, and why these fractions describe two recorded days and not future performance.

Notebooksthat check your answers.

Open them in Google Colab. Each answer is checked as you go: correct, wrong with the expected value, or not answered yet.

  • Probability basicsLessons 3.1 Courses plan
  • Conditional probability, independence and Bayes' theoremLessons 3.2 and 3.3 Courses plan
  • Random variables, expectation and varianceLessons 3.4 Courses plan
  • Stats 3.5 (optional) · Monte Carlo simulationLessons 3.5 Courses plan
  • Stats Module 3 project · A fraud review-queue briefThe module project, with checks Courses plan

Referencefor revising and for interviews.

The cheat sheet is one page of the module’s terms, rules and gotchas. The interview questions come with model answers.