All stories
AI learning

Python for machine learning: what to learn first, and what to skip

The Python you need for machine learning, in order: core language, files and testing, then NumPy, pandas, Matplotlib and scikit-learn, plus what you can safely leave for later.

To learn Python for machine learning, start with the core language: variables and types, if and loops, lists and dictionaries, functions, and reading a file. Then learn four libraries in this order: NumPy for arrays, pandas for tables, Matplotlib for charts and scikit-learn for models. You can skip web frameworks, GUI toolkits, metaclasses, async code and most of the standard library at first. Once you can load a CSV, clean it, plot it and fit a baseline model, you know enough Python to move on to the machine learning itself.

That is the short answer. The rest of this post explains what each stage gives you, what "enough" looks like, and which topics are worth postponing.

Why is Python used for AI and ML?

Python is not the fastest language, but it is readable, it runs one statement at a time so you see results immediately, and the libraries machine learning depends on are written for it: NumPy, pandas, Matplotlib, scikit-learn, and for deep learning PyTorch and TensorFlow. The heavy number crunching inside those libraries runs as compiled code, so your Python mostly describes what to compute, and the library does the work quickly.

That is also why the Python you need for AI and ML is narrower than "all of Python". You need to read and write the code that sits around those libraries: loading data, reshaping it, passing it to a model and checking the result. The lesson Why Python? covers this in a short free video, including where code is written and run (Jupyter notebooks and Google Colab).

What to learn first: the core language

Learn these before you open a machine learning library. Each one shows up in almost every notebook you will read.

TopicWhy machine learning code needs itLesson
Variables, types and operatorsEvery value is an object with a type; / and // behave differentlyData types, Operators
if, for and rangeLooping over rows, files, epochs and settingsfor loops and range
Lists and dictionariesCollections of samples; counting and groupingLists, Dictionaries
StringsCleaning and splitting text is the first step of every text modelStrings
ComprehensionsBuilding a list or dict in one readable lineComprehensions
Functions and argumentsLibrary calls are full of keyword arguments and defaultsDefining functions, Arguments and scope
Lambdas and functions as valuessorted(key=...), apply and callbacksLambdas

If you are completely new, begin with Run your first Python code, which explains what happens when a notebook cell runs, including the kernel and execution order. Notebooks that "worked yesterday" usually fail because cells ran in a different order, so this is worth understanding early.

A good checkpoint for this stage: read a text file, count how often each word appears using a dictionary, and print the ten most common words. If you can write that without copying it, you are ready for the next stage.

What to learn next: files, environments and tests

This is the stage most self-taught learners skip, and the one that makes later projects reproducible.

  • Modules and environments. Know how import works and how to install packages into a project without breaking another one. The lesson on uv, packages and environments covers uv, pyproject.toml and what a virtual environment actually is.
  • Files. Read and write text, CSV and JSON, with paths built by pathlib. See Working with files.
  • Errors and debugging. Read a traceback from the bottom up, handle errors deliberately with exceptions, and find bugs with a debugger rather than guesswork.
  • Tests. A short pytest test for a data-cleaning function catches mistakes before they reach a model. See Testing with assert and pytest.

Classes: learn just enough to read scikit-learn

You do not need to design class hierarchies to do machine learning. You do need to read code like model = LogisticRegression(C=0.5) followed by model.fit(X, y) and model.coef_. Learn classes and objects well enough to understand __init__, self, attributes and methods.

Then take the estimator pattern. It has you write a tiny model class with fit and predict, storing what it learned in attributes with a trailing underscore. That is the exact shape every scikit-learn model has, and once you have written one yourself the library stops feeling like magic.

The four libraries, in order

The free Data Tools track takes these in the order they build on each other.

  1. NumPy. Arrays are faster and smaller than lists because they hold one typed block of memory. Learn shape, dtype and reshaping, indexing, slicing and boolean masks, and vectorized operations and broadcasting. Broadcasting is where most silent shape bugs come from, so practice predicting the result shape before you run the code.
  2. pandas. A DataFrame is a table with named columns and an index. Learn selecting and filtering, cleaning data (missing values, types and duplicates) and groupby and aggregation.
  3. Matplotlib. Look at data before you model it. Meet Matplotlib covers a first labeled plot and when a line or a scatter answers the question.
  4. scikit-learn. Meet scikit-learn fits a first model against a baseline, and Estimators: fit, predict, transform explains the three verbs every model shares.

Here is a small example that uses the first two together. It runs as written with NumPy and pandas installed.

import numpy as np
import pandas as pd

df = pd.DataFrame({
    "city": ["Austin", "Boston", "Austin", "Denver", "Boston"],
    "sqft": [1400, 900, 2100, None, 1200],
    "price": [410_000, 520_000, 615_000, 480_000, 650_000],
})

df["sqft"] = df["sqft"].fillna(df["sqft"].median())  # clean a missing value
print(df.groupby("city")["price"].mean())             # a statistic per group

X = df[["sqft", "price"]].to_numpy(dtype=float)
X_std = (X - X.mean(axis=0)) / X.std(axis=0)          # broadcasting, one line
print(X_std.shape, X_std.mean(axis=0).round(6))

The last two lines standardize each column so it has mean 0 and standard deviation 1. X.mean(axis=0) has shape (2,), and NumPy broadcasts it across all five rows. If you can explain why that works, and what would go wrong with axis=1, you understand the most important idea in the NumPy section.

What to skip (for now)

None of these is useless. They just do not stand between you and your first model.

  • Web frameworks such as Django or Flask. Useful later for serving a model; not needed to train one.
  • GUI toolkits, game libraries and scraping frameworks.
  • Advanced language features: metaclasses, descriptors, async/await, writing your own decorators and context managers.
  • Memorizing the standard library. Learn pathlib, json, csv, collections and random; look the rest up when you need it.
  • Deep learning frameworks before you have trained a classical model. PyTorch makes more sense once you know what a loss, a gradient and a test set are.
  • Exotic pandas tricks. merge, pivot_table and time series resampling matter, but you can learn them the first time a project needs them.

One topic people often skip that is worth a short detour: time and space complexity. Knowing why x in my_list is slow and x in my_set is fast will save you from loops that take minutes instead of seconds.

How much Python do I need before machine learning?

Enough to write a script that loads a CSV, cleans it, computes a statistic per group, plots one chart, and fits a scikit-learn model against a baseline, without looking up every line. You will keep learning Python as you go; you do not need to finish a Python course before you touch a model.

Python is one of three prerequisites. The other two are some math and some statistics, covered in prerequisites for machine learning. When you are ready to see the whole workflow in code, the machine learning code in Python post builds a first model in about 30 lines, and where to start with machine learning puts all of it in order.

Your next step

Open the Python track and start with lesson 1.1. Every lesson's parts and key terms are free to read, and the first two lessons have free videos. A free account adds the quizzes and saves your progress, with no payment. When the core language feels comfortable, move on to the Data Tools track.

Learn machine learning with Python: interview questions with model answers, real-world applications, practice and projects.Start learning free

Directory listings

  • SchoolWhool on Siteefy