Machine learning code in Python: your first model in about 30 lines
A complete, runnable scikit-learn example: load a built-in dataset, split it, beat a baseline, fit a model and evaluate it, with every line explained.
Here is complete machine learning code in Python: it loads scikit-learn's built-in breast cancer dataset, holds back a quarter of it as a test set, scores a do-nothing baseline, trains a logistic regression, and evaluates it on the rows it never saw. It is 31 lines including comments and blank lines, and it runs as written once scikit-learn is installed. Every serious project follows the same five steps, just with messier data: load, split, baseline, fit, evaluate.
Copy the code first, run it, then read the line-by-line explanation below.
The complete code
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.dummy import DummyClassifier
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, confusion_matrix, classification_report
# 1. Load a built-in dataset: 569 tumors, 30 numeric features each.
X, y = load_breast_cancer(return_X_y=True)
print("X shape:", X.shape, " y shape:", y.shape)
# 2. Hold back 25% of the rows as a test set the model never sees in training.
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=42, stratify=y
)
# 3. A baseline: always predict the most common class.
baseline = DummyClassifier(strategy="most_frequent")
baseline.fit(X_train, y_train)
print("Baseline accuracy:", round(baseline.score(X_test, y_test), 3))
# 4. The model: scale the features, then fit a logistic regression.
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_train, y_train)
# 5. Predict on the test set and evaluate.
y_pred = model.predict(X_test)
print("Model accuracy:", round(accuracy_score(y_test, y_pred), 3))
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred, target_names=["malignant", "benign"]))
How do I run this Python code for machine learning?
You need Python 3 and one package. In a terminal:
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
pip install scikit-learn
python first_model.py
scikit-learn installs NumPy and SciPy for you. In Google Colab, scikit-learn is already installed, so you can paste the code into a cell and run it. If environments are new to you, the lesson on uv, packages and environments explains what a virtual environment is and why each project gets its own.
What the output looks like
With scikit-learn 1.9 on Python 3.14, the script printed:
X shape: (569, 30) y shape: (569,)
Baseline accuracy: 0.629
Model accuracy: 0.986
[[52 1]
[ 1 89]]
precision recall f1-score support
malignant 0.98 0.98 0.98 53
benign 0.99 0.99 0.99 90
accuracy 0.99 143
macro avg 0.99 0.99 0.99 143
weighted avg 0.99 0.99 0.99 143
Because the split uses a fixed random_state, you should see the same numbers or numbers very close to them. A different library version can change the last digit.
Line by line: loading the data
The imports. Each import names one tool. scikit-learn keeps datasets, splitting, models, preprocessing and metrics in separate modules, so the import list doubles as a table of contents for the script.
X, y = load_breast_cancer(return_X_y=True) loads a small, clean medical dataset that ships with scikit-learn. X is a NumPy array with one row per tumor and one column per measurement (radius, texture and so on). y holds the label for each row: 0 for malignant, 1 for benign. This X and y split, a matrix of features and a vector of labels, is the input format of almost every scikit-learn model.
The print of shapes is a habit worth keeping. (569, 30) tells you there are 569 samples and 30 features, and that y has one label per sample. Most beginner bugs are shape bugs; the shape, dtype and reshaping lesson covers how to read them.
Line by line: the train/test split
train_test_split(...) shuffles the rows and returns four arrays. The model learns only from X_train and y_train; X_test and y_test are kept back to measure it. Without this step, you would be grading the model on questions it had already seen, and the score would mean very little.
test_size=0.25keeps 25% of the rows (143 of 569) for testing.random_state=42fixes the shuffle so the result is reproducible. The number itself does not matter.stratify=ykeeps the proportion of malignant and benign cases the same in both parts.
Line by line: the baseline
DummyClassifier(strategy="most_frequent") is a model that ignores the features and always predicts the most common class. It scores 0.629, because about 63% of the test rows are benign. That number is the bar a real model has to clear. An accuracy of 0.63 sounds respectable until you see that predicting "benign" every time gets the same.
Notice that the baseline uses the same two verbs as every other model: fit to learn (here, which class is most common) and score to evaluate. That shared interface is the main idea of scikit-learn, and it is what the lessons Meet scikit-learn and Estimators: fit, predict, transform teach.
Line by line: the model
make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000)) chains two steps into one object.
StandardScalerrescales each feature to mean 0 and standard deviation 1. The raw features are on very different scales (some are below 1, others run into the hundreds or thousands), and logistic regression trains faster and more reliably when they are comparable.LogisticRegressionis a linear classifier: it computes a weighted sum of the features and turns it into a probability. Despite the name, it is used for classification, not regression.max_iter=1000gives the solver enough iterations to converge, so you do not get a convergence warning.
model.fit(X_train, y_train) is where learning happens. The scaler learns each feature's mean and standard deviation from the training rows only, and the classifier learns one weight per feature. Putting the scaler inside the pipeline matters: if you scaled the whole dataset before splitting, the test rows would leak into the training statistics. The lesson on pipelines shows this pattern with cross-validation.
Line by line: evaluating the model
y_pred = model.predict(X_test) produces a predicted label for each of the 143 test rows. Calling predict on the pipeline scales the test rows with the training statistics, then classifies them.
accuracy_score is the fraction of predictions that were right: 0.986, against the baseline's 0.629.
confusion_matrix shows which mistakes were made. Rows are the true classes and columns the predicted ones, in label order (malignant, then benign). Here 52 malignant tumors were caught and 1 was missed, and 89 benign tumors were correct with 1 flagged as malignant. In a medical setting, the missed malignant case matters far more than the false alarm, which is why accuracy alone is never the whole story.
classification_report gives precision (of the cases predicted as a class, how many were right) and recall (of the cases truly in a class, how many were found) for each class.
Is 98.6% accuracy good?
On this dataset, it is a strong result for a simple model, but be careful what you conclude. The dataset is small, clean and well studied, which real data rarely is. The test set has 143 rows, so a single different prediction moves accuracy by about 0.7 percentage points. A confidence interval is the honest way to report a score from a sample this size, and cross-validation gives a steadier estimate than one split.
What to change next
Once this runs, try one change at a time and predict the result before you run it:
- Swap
LogisticRegression(max_iter=1000)forKNeighborsClassifier()(fromsklearn.neighbors). The rest of the code stays the same. - Remove
StandardScaler()from the pipeline and see what happens to k-nearest neighbors. - Replace the single split with
cross_val_score(model, X, y, cv=5)fromsklearn.model_selection. - Swap the dataset for
load_iris, which has three classes instead of two.
Each experiment teaches one idea: the shared estimator interface, why scaling matters for distance-based models, and why one split is a noisy measurement.
What this code does not teach you
Thirty lines of scikit-learn calls are a good start, but calling fit is not the same as understanding it. To know what fit actually does inside a linear model, learn how linear regression works and how gradient descent finds its weights. Both are free video lessons in the Machine Learning course. If the Python here felt unfamiliar, start with Python for machine learning; if the math did, read prerequisites for machine learning.
Your next step
Work through the scikit-learn module of the Data Tools track, which builds this same workflow step by step. Lesson parts and key terms are free to read, and a free account adds the quizzes and saves your progress, with no payment. For the order to learn everything else in, see where to start with machine learning.