All stories
AI learning

Prerequisites for machine learning: the math, statistics and Python you need

The real prerequisites for machine learning: working Python, a short list of linear algebra and calculus ideas, and the statistics that tell you whether a model works. With a checklist to test yourself.

The prerequisites for machine learning are three skills: Python you can write without copying, a small amount of math (vectors, matrices, derivatives and gradients, logarithms), and introductory statistics and probability (distributions, spread, conditional probability, and how to tell a real difference from noise). You do not need a math degree, and you do not need to learn every topic before you start. You need each idea well enough to recognize it when a model uses it, and you can learn much of the math alongside your first models.

Below is what each prerequisite covers, why machine learning needs it, and a checklist you can use to decide whether you are ready.

What are the skills needed for machine learning?

SkillWhat you actually needWhy a model needs it
PythonFunctions, lists, dictionaries, files, classes at a reading levelEvery library, dataset and experiment is driven from Python
Data toolsNumPy arrays, pandas tables, a basic chart, the scikit-learn APIData has to be loaded, cleaned and shaped before any model sees it
Linear algebraVectors, the dot product, matrices, the matrix–vector productA dataset is a matrix; a linear model's predictions are one product
CalculusDerivatives, the chain rule, partial derivatives, the gradientTraining lowers a loss by following its gradient
StatisticsMean, median, variance, distributions, correlation, samplingYou need to understand the data before you trust a model built on it
ProbabilityConditional probability, Bayes' theorem, expectation, likelihoodClassifiers output probabilities; log loss comes from likelihood

Notice what is missing: proofs, integration techniques, abstract algebra, and most of a typical statistics syllabus. Those can matter for research, but they are not prerequisites for building and evaluating classical models.

Python: the prerequisite you cannot skip

Math can be learned in parallel with your first model. Python cannot, because every exercise is written in it. You should be able to write a function, loop over a list, use a dictionary to count things, read a CSV file and handle an error. You should also be able to read a class well enough to understand model.fit(X, y) and model.coef_.

The free Python track is built for this: it teaches the language the way machine learning code uses it, from running your first code to the estimator pattern, where you write a tiny model class with fit and predict. For what to prioritize and what to skip, see Python for machine learning.

After the language, learn the four libraries in the Data Tools track: NumPy, pandas, Matplotlib and scikit-learn. Linear algebra in NumPy is a good bridge to the next section, because it writes the matrix math of machine learning directly as code.

How much math do you need?

Less than most people fear, but more than none. The SchoolWhool Maths track covers the set machine learning uses, each idea shown on a picture before the formula:

  • Functions, logarithms and notation. Read ŷᵢ = Σⱼ wⱼ xᵢⱼ aloud in words. Know that a log turns a product into a sum, which is why likelihoods are always computed as log-likelihoods.
  • Vectors. A row of data is a point and an arrow. The dot product measures alignment, and a weighted sum of features (a linear model's prediction) is a dot product.
  • Matrices. A dataset is an n × d matrix. Every prediction at once is the product Xw. The transpose, the inverse, and later eigenvectors and the SVD (behind PCA) all build on this.
  • Calculus. A derivative is a rate of change. A gradient points uphill, so stepping against it lowers the loss. That one sentence is gradient descent. The chain rule is how the gradient of a long computation is calculated, which is the core of training neural networks.

The Maths track is being recorded; its first lesson page, Derivatives: rate of change and the tangent, has a free video. The other lessons appear in the track outline and open as they are published.

Here is most of that math in ten lines of NumPy. It runs as written.

import numpy as np

X = np.array([[1.0, 2.0], [2.0, 0.5], [3.0, 1.5]])  # 3 samples, 2 features
y = np.array([5.0, 4.0, 7.5])                      # the targets
w = np.zeros(2)                                    # the weights, before training

for step in range(500):
    y_hat = X @ w                                  # every prediction: dot products
    error = y_hat - y
    loss = (error ** 2).mean()                     # mean squared error
    grad = 2 * X.T @ error / len(y)                # the gradient of the loss
    w -= 0.05 * grad                               # one step of gradient descent

print("weights:", w.round(3), " loss:", round(loss, 4))

It prints weights of about [1.622 1.705] and a loss close to zero. If you can explain each commented line, you already have the linear algebra and calculus that the first half of a machine learning course uses. If you cannot yet, that is the list of things to learn. The free gradient descent lesson in the Machine Learning course walks through the idea with pictures.

What statistics and probability do you need?

Statistics is the prerequisite people underrate. A model is only as good as the data it learned from, and reading a test score correctly is a statistical question. The Statistics track covers this, and these lessons are the core:

Hypothesis testing and information theory (entropy, cross-entropy and log loss) matter too, but they can wait until you meet them in a model.

Do I need all of this before I start?

No. A practical order is: Python first, then the data tools, then start machine learning while you learn the math and statistics alongside it. Learn the derivative right before you meet gradient descent, the dot product right before linear regression, and Bayes' theorem right before Naive Bayes. Ideas stick better when they solve a problem you have just met.

The one thing to avoid is the opposite extreme: spending months on abstract math without ever training a model. The where to start with machine learning post sets out a four-step order that interleaves the two.

A readiness checklist

You are ready to start a first machine learning course if you can do most of these:

  • Write a Python function that takes a list and returns a dictionary of counts.
  • Load a CSV with pandas, fill or drop missing values, and compute an average per group.
  • Explain what X.shape == (569, 30) means for a dataset.
  • Compute a dot product by hand for two short vectors.
  • Say what the sign of a derivative tells you about a curve.
  • Explain why the median can be a better "typical value" than the mean.
  • Explain why correlation is not causation, with an example.
  • Explain why a model must be scored on data it did not train on.

If two or three are shaky, start anyway and fill the gaps as they come up. If most are shaky, spend a few weeks on the Python and Statistics tracks first.

Your next step

Start with whichever prerequisite is weakest: the Python track, the Statistics track or the Maths track. Lesson parts and key terms are free to read, and so are the free videos. A free account adds the quizzes and saves your progress, with no payment. When you want to see a whole model in code, read machine learning code in Python.

Learn machine learning with Python: interview questions with model answers, real-world applications, practice and projects.Start learning free

Directory listings

  • SchoolWhool on Siteefy