Stats 5.4Concept6 parts

The Bootstrap

Get an interval for any statistic by resampling the data you have.

In this lesson6 parts

  1. 01The comparison we want to trust
  2. 02Compute the observed paired score
  3. 03Build one paired bootstrap replicate
  4. 04Build intervals from the bootstrap distributions
  5. 05Break and restore pairing
  6. 06Reproduce and challenge

Key terms

The words this lesson introduces, each in one line. The module’s glossary collects them all.

paired comparison
Two results measured on the same observational units.
ROC-AUC
Rank-based score comparing positive and negative labels: the share of positive–negative pairs where the positive scores higher, ties counting one half.
validation score
Performance calculated on data withheld from fitting.
observed difference
B's score minus A's score on the original validation set.
bootstrap (narration)
Drawing new samples, with replacement, from the one observed set itself.
bootstrap replicate
One resampled dataset, drawn with replacement from the observed rows.
paired resampling
Applying one draw of unit indices to every measurement on those units.
bootstrap distribution
The collection of a statistic from many bootstrap replicates.
percentile bootstrap interval
Bounds read from specified quantiles of a bootstrap distribution.
fixed-model uncertainty
Variation from resampled validation rows while fitted models stay unchanged.

Quiz 5 questions

Your first pick on each question is the one that counts, and a right one earns a coin. Getting one wrong here is how the lesson sticks.

Practice

Problems to solve in your own notebook. Each states the problem, not the steps: working out the steps is the exercise. Level A applies the lesson, B combines it with earlier ones, C stretches it.

The self-checking notebook for this lesson is The bootstrap.

Common mistakes

What you will see when it goes wrong, why it happens, and the fix.

Where it’s used

Where this lesson’s ideas turn up in real work.