All stories
AI learning

Statistics for data science and machine learning: what to learn, in order

The statistics a data scientist or machine learning engineer actually uses: describing data, probability, distributions, estimation, testing and the information theory behind loss functions. A seven-step order with what to skip.

Data science needs less statistics than a full degree teaches, but it needs the right parts, in an order where each one builds on the last. In practice that is seven topics: describing data, reading evidence critically, probability, distributions, estimation with uncertainty, hypothesis testing, and the information theory behind classification losses. You can learn the first four before your first model and the rest alongside it.

Here is what each step covers, why it matters in real data work, and the lessons in the free Statistics track that teach it.

The short answer

StepTopicYou can do this afterwards
1Describing dataSummarize any column and pick the right chart for it
2Evidence and EDATell a real pattern from a misleading one before modeling
3ProbabilityReason about rare events, classifiers and conditional evidence
4DistributionsRecognize the shape of data and what a model assumes about it
5EstimationPut an honest interval around any number, including a model score
6Hypothesis testingDecide whether a difference or an A/B result is more than noise
7Information theoryUnderstand log loss, cross-entropy and how trees choose splits

1. Describing data

Every analysis starts by summarizing what is there. You need a typical value, a measure of spread and a picture of the shape:

In a model this is feature inspection: spotting skew, outliers, impossible values and features that duplicate each other.

2. Evidence and exploratory data analysis

The most expensive statistical mistakes come from trusting a pattern that the data cannot support. This step is about reading evidence critically: how charts can mislead, missing denominators and base rates, Simpson's paradox and survivorship bias. The post on misleading statistics walks through seven of these with examples.

It also treats EDA as a sequence of decisions, not a pile of charts, ending in a one-page data report for a model.

3. Probability

Classifiers output probabilities, and most of the reasoning around them is conditional probability. Learn probability basics, conditional probability and independence and Bayes' theorem, which explains why a 99% accurate test for a rare condition is wrong most of the times it says yes. Then expected value and variance, which is how you compare two decision rules with different costs, and Monte Carlo simulation, which answers probability questions in code when the algebra is hard.

4. Distributions

A distribution is the shape of the process that produced your data. The ones that matter most in practice:

You will also use the Q-Q plot to check whether a model's residuals look the way it assumes.

5. Estimation and uncertainty

Every number computed from a sample, including a model's accuracy on a test set, is an estimate. This step is about how much it could move:

6. Hypothesis testing

Testing answers one question: could this difference be noise? Learn hypothesis testing and the p-value through a simulation first, then permutation tests, the t-test and the chi-square test, and how tests go wrong through multiple comparisons and repeated looks. The step ends with designing and reading an A/B test for a deployed model, which is one of the most common statistics tasks in industry.

7. Information theory for machine learning

Classification models are trained by minimizing cross-entropy, also called log loss. Decision trees choose splits by entropy or Gini impurity. Model monitoring compares distributions with KL divergence, and feature selection can use mutual information. These are short lessons once probability is in place, and they make the loss functions you meet in machine learning readable rather than magic.

What you can skip for now

For applied data science and classical machine learning you can skip, at first: measure-theoretic probability, proofs of the central limit theorem, most named distributions beyond the ones above, ANOVA tables, and hand calculation from statistical tables. Software computes all of these. What it cannot do is tell you which comparison is fair, whether a sample represents the population, or whether a difference matters, and that judgment is what the steps above build.

Do you need all of it before your first model?

No. Steps 1 to 4 are enough to start your first models, and you can learn them alongside Python. Steps 5 to 7 are easier to learn while you are already training models, because each one answers a question you will have just met: how sure am I of this score, is this improvement real, and what is this loss function measuring.

For where statistics sits among the other prerequisites, see prerequisites for machine learning. For the full learning order across Python, data tools, math and statistics, see where to start with machine learning.

Your next step

Start the Statistics track at the first module, why statistics matters. Lesson parts and key terms are free to read. A free account adds the quizzes and saves your progress, with no payment. If you are learning machine learning seriously, the learner community below is where course access for selected learners is shared.

Serious about learning machine learning? Join the learner community: selected learners get free course access in exchange for honest feedback.
Start learning free

Directory listings

  • SchoolWhool on Siteefy