Python for AI

Python for machine learning

The bridge lesson. You do the same small job twice, once in plain Python and once with NumPy and pandas, and see exactly which problem those libraries were built to solve.

Read these first

On this page 7
  1. Why you should care
  2. The five steps, every time
  3. Why the libraries exist, in one sentence
  4. A real example you have seen
  5. An honest word about where the time goes
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Machine learning code is ordinary Python, arranged in the same five steps every single time.

Think of a bank cashier counting notes. Twenty notes, and hands are fine. Two thousand notes, and hands are hopeless. So the cashier reaches for a counting machine. The notes also get stacked into neat bundles first, so the machine can take them.

Nothing about counting changed. The scale changed, so the tools and the arrangement had to change with it.

That is exactly the story of NumPy and pandas. Plain Python counts by hand. Those two libraries are the machine and the neat bundles.

Why you should care

You have finished the language. You know names, containers, loops and functions. You have met arrays, tables and charts.

The natural question is: what do I actually do with all of it? This lesson answers that, and then hands you over to the machine learning section.

The five steps, every time

Whether the model is a simple line-fitter or a network with a billion knobs, the script around it has the same skeleton.

   1. get the data      →   a file, a database, an export from somebody
   2. clean it          →   fix wrong kinds of value, fill or drop the holes
   3. split it          →   hide some rows so you can test honestly later
   4. teach the model   →   show it the rows it is allowed to see
   5. check it          →   score it on the hidden rows

Once you recognise this shape, unfamiliar AI code becomes readable. You stop reading it line by line and start asking which of the five steps you are looking at.

Why the libraries exist, in one sentence

Steps two and three involve touching every value in the dataset.

With plain Python you write a loop, and the loop runs one value at a time. With NumPy you describe the change once and it happens to the whole pile. With pandas you do the same thing to a whole named column of a table.

The Developer tab does the identical job both ways, so you can see the difference rather than take it on trust.

A real example you have seen

Your bank sends a message when a payment looks unusual. Behind it, a model was shown millions of past payments labelled as fine or as fraud.

Somebody wrote step one: pull the payments. Step two: fix the dates, drop the broken rows. Step three: hold back last month. Step four: train. Step five: check how many frauds it caught and how many honest payments it wrongly blocked.

That last number is why step five is not optional. A model that flags everything catches all fraud and is useless.

An honest word about where the time goes

Beginners expect the model to be the hard part. It rarely is.

The model is often five lines, because somebody else wrote the library. Steps one and two are where the weeks go. You chase a column with three different spellings of the same city. A date stored as text. A price with a currency symbol glued to it.

This is not a failure of your skills. It is the job. Practitioners across the industry report the same split, and the people who are good at it are good at the boring part.

The second honest thing: finishing this section does not mean you know machine learning. It means you can now read and write the code that machine learning is expressed in. The ideas start in the next section.

Remember this

  • Every ML script is the same five steps: get, clean, split, teach, check.
  • NumPy and pandas exist because steps two and three touch every value, and plain Python touches them one at a time.
  • Most of the work is cleaning data, and that is normal rather than a sign you are doing it wrong.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas

The first half of this page imports nothing at all. Run it on any Python you have.

Everything here fits in a few kilobytes of memory and finishes instantly on a CPU. No GPU, no download, no dataset.

The job

Five house records. Prices in lakhs. The data arrives the way real data arrives: every value is text, and one value is missing entirely.

The task is the front end of every supervised learning script. Convert the text to numbers, fill the hole, put every column on the same scale, apply a set of weights, and measure how wrong the result is.

Doing it with plain Python

plain_python.py
rows = [
    {"area": "650",  "rooms": "2", "age": "10", "price": "35"},
    {"area": "800",  "rooms": "3", "age": "5",  "price": "48"},
    {"area": "1200", "rooms": "3", "age": "20", "price": "60"},
    {"area": "950",  "rooms": "",  "age": "8",  "price": "52"},   # rooms is missing
    {"area": "1500", "rooms": "4", "age": "2",  "price": "85"},
]

FEATURES = ["area", "rooms", "age"]

# step 2a: text -> numbers, with the blank becoming an explicit None
clean = []
for row in rows:
    clean.append({k: (float(v) if v != "" else None) for k, v in row.items()})

# step 2b: fill each hole with the average of the values that are present
for name in FEATURES:
    known = [r[name] for r in clean if r[name] is not None]
    fill = sum(known) / len(known)
    for r in clean:
        if r[name] is None:
            r[name] = fill

# step 2c: standardise, so area in hundreds does not drown rooms in single digits
stats = {}
for name in FEATURES:
    col = [r[name] for r in clean]
    mean = sum(col) / len(col)
    var = sum((x - mean) ** 2 for x in col) / len(col)
    stats[name] = (mean, var ** 0.5)

X = []
for r in clean:
    X.append([(r[n] - stats[n][0]) / stats[n][1] for n in FEATURES])

y = [r["price"] for r in clean]

# step 4: weights typed by hand, so this file prints the same numbers for you
w = [0.9, 0.3, -0.2]
bias = 56.0
preds = []
for xrow in X:
    total = bias
    for value, weight in zip(xrow, w):
        total = total + value * weight
    preds.append(total)

# step 5: how far off are we, on average
errors = [abs(p - t) for p, t in zip(preds, y)]

print("standardised first row:", [round(v, 3) for v in X[0]])
print("predictions:", [round(p, 2) for p in preds])
print("mean absolute error:", round(sum(errors) / len(errors), 3))
Output
standardised first row: [-1.229, -1.581, 0.163]
predictions: [54.39, 55.47, 56.18, 55.82, 58.14]
mean absolute error: 12.273

That works. It is also 29 lines of logic, and every one of them is a place to make a mistake.

Count the traps you had to avoid. The fill average must skip the missing entries or it crashes on None. The standardising loop must run over columns while the prediction loop runs over rows, and nothing in the code stops you mixing them up. zip(xrow, w) silently truncates if the weight list is the wrong length. And the fill step divides by len(known), which is zero if a column is entirely empty.

Now scale it in your head. Five rows and three columns became fifty thousand rows and two hundred columns. The code above does not change shape at all. It stops finishing, and that is the only difference.

The same job with pandas and NumPy

with_numpy.py
import numpy as np
import pandas as pd

df = pd.DataFrame([
    {"area": "650",  "rooms": "2", "age": "10", "price": "35"},
    {"area": "800",  "rooms": "3", "age": "5",  "price": "48"},
    {"area": "1200", "rooms": "3", "age": "20", "price": "60"},
    {"area": "950",  "rooms": "",  "age": "8",  "price": "52"},
    {"area": "1500", "rooms": "4", "age": "2",  "price": "85"},
])

df = df.apply(pd.to_numeric, errors="coerce")   # text -> numbers, "" becomes NaN
df = df.fillna(df.mean())                       # every hole gets its column average

X = df[["area", "rooms", "age"]].to_numpy()
y = df["price"].to_numpy()

X = (X - X.mean(axis=0)) / X.std(axis=0)        # standardise every column at once
preds = X @ np.array([0.9, 0.3, -0.2]) + 56.0
errors = np.abs(preds - y)

print("standardised first row:", [round(float(v), 3) for v in X[0]])
print("predictions:", [round(float(p), 2) for p in preds])
print("mean absolute error:", round(float(errors.mean()), 3))
Output
standardised first row: [-1.229, -1.581, 0.163]
predictions: [54.39, 55.47, 56.18, 55.82, 58.14]
mean absolute error: 12.273

Identical numbers. Seven lines of logic instead of twenty-nine.

That file runs clean and prints those exact numbers on pandas 2.2 and on pandas 3.0. pandas moves quickly, so a later version may add a FutureWarning about one of those calls. The numbers are unaffected; read the warning and follow what it suggests.

What actually changed

The line count is the least interesting part. Three deeper things happened.

The loops moved out of Python. X.mean(axis=0) is one instruction from Python's point of view, and a tight compiled loop underneath. You measured that gap yourself in the NumPy lesson; this is the same effect applied to a whole pipeline.

The bookkeeping moved into the data structure. In the plain version, stats had to be a dictionary you built and indexed correctly. In the second version, the column-to-statistic mapping is implied by the array's shape. Code you do not write cannot be wrong.

The missing value became a first-class thing. NaN travels through arithmetic and stays visible. A hand-rolled None raises TypeError the moment it reaches a subtraction, usually three functions away from where it entered.

Memory, the reason that is easy to miss

memory.py
import sys
import numpy as np

print("one Python float object:", sys.getsizeof(0.0), "bytes")   # 64-bit CPython

rows, cols = 50_000, 20
arr = np.zeros((rows, cols), dtype=np.float64)
print("one number inside an array:", arr.itemsize, "bytes")
print("whole table as an array   :", arr.nbytes / 1e6, "MB")
Output
one Python float object: 24 bytes
one number inside an array: 8 bytes
whole table as an array   : 8.0 MB

The exact byte count from sys.getsizeof can shift between CPython releases, so treat 24 as "roughly three times 8" rather than as a constant.

Three times is the optimistic reading. The same million numbers held as Python objects also need a pointer each, plus the dictionaries that group them into rows, and they sit scattered across memory so the processor's cache never helps. The array is one unbroken block that a compiled kernel can stream through.

The Python you actually need for ML

You have covered all of it. Here is the honest division, so you can stop worrying about the rest of the language for now.

Use constantlyCan wait
Variables, ints, floats, strings, bools, NoneClasses and inheritance
Lists, dictionaries, tuples, setsDecorators, async, threading
for loops, if/elif/elseMetaclasses, descriptors, operator overloading
Functions, default arguments, return valuesRegular expressions, until you handle raw text
List comprehensionsPackaging, virtual environment plumbing
import, and reading library documentationEverything else

You will meet classes when you write a PyTorch model, and decorators when you cache something. Both are one afternoon each, later, with a reason to learn them.

Common mistakes

1. Two libraries, two definitions of standard deviation.

ddof.py
import numpy as np
import pandas as pd

x = [2, 4, 4, 4, 5, 5, 7, 9]
print("numpy  std:", np.std(x))                      # divides by n
print("pandas std:", round(pd.Series(x).std(), 6))   # divides by n - 1
print("matched   :", round(pd.Series(x).std(ddof=0), 6))
Output
numpy  std: 2.0
pandas std: 2.13809
matched   : 2.0

Neither is wrong. NumPy defaults to the population formula, pandas to the sample formula. Mixing them inside one pipeline produces features that are quietly scaled differently in training and in serving. Pick one, pass ddof explicitly, and write it down.

2. Learning the scaling numbers from the test rows.

split.py
import numpy as np

X = np.array([[10.0], [20.0], [30.0], [40.0], [50.0],
              [60.0], [70.0], [80.0], [90.0], [100.0]])

X_train, X_test = X[:8], X[8:]      # a fixed split, so this page is reproducible

mean = X_train.mean(axis=0)         # learned from TRAIN only
sd = X_train.std(axis=0)
print("train mean:", mean, "train sd:", np.round(sd, 4))

X_test_s = (X_test - mean) / sd     # the same two numbers, applied to test
print("scaled test rows:", [round(float(v), 3) for v in X_test_s.ravel()])

print("mean if you leak the test rows:", X.mean(axis=0))
Output
train mean: [45.] train sd: [22.9129]
scaled test rows: [1.964, 2.4]
mean if you leak the test rows: [55.]

Scaling with 55 instead of 45 lets information from the test rows reach the model. This is data leakage, and it makes your test score better than the real world will be. The bug is invisible: nothing crashes, and the number on your screen goes up. See train, test and validation splits.

3. A one-dimensional array where a two-dimensional one is expected.

python
import numpy as np

X = np.array([1.0, 2.0, 3.0])
print("shape:", X.shape)
print("as a column:", X.reshape(-1, 1).shape)
Output
shape: (3,)
as a column: (3, 1)

Model libraries want features shaped (n_samples, n_features) and targets shaped (n_samples,). Hand a flat array where a table is expected and scikit-learn answers with Expected 2D array, got 1D array instead. The -1 in reshape means "work this dimension out from the total". Print .shape before every call into a library and most of these disappear.

4. Results you cannot reproduce.

seed.py
import numpy as np

a = np.random.default_rng(42).normal(size=4)
b = np.random.default_rng(42).normal(size=4)
c = np.random.default_rng(7).normal(size=4)
print("same seed  ->", np.allclose(a, b))
print("other seed ->", np.allclose(a, c))
Output
same seed  -> True
other seed -> False

Shuffling, splitting and weight initialisation all draw random numbers. Without a fixed seed, two runs of the same script give two different scores, and you cannot tell an improvement from noise. Create an explicit generator and pass it around rather than calling np.random.seed, which mutates hidden global state that any library can also change.

5. Judging a model by the number that flatters it.

The mean absolute error above is 12.273. Is that good? On its own the number cannot tell you, because there is no baseline.

Predict the average price for every house and the error is 13.2. So those hand-typed weights beat the dumbest possible model by less than one lakh, which is close to nothing. A number that looked like a result was almost entirely the baseline.

Always compute the trivial baseline and report your score next to it. A model that cannot beat the average is not a model.

Try it yourself

Take with_numpy.py and add a sixth house whose age is blank as well as a seventh with a valid row. Confirm the fill step handles two missing values in different columns without any code change. Then do the same to plain_python.py and count the lines you had to touch.

Next, replace the hand-typed weights with weights found by least squares: w, *_ = np.linalg.lstsq(np.c_[X, np.ones(len(X))], y, rcond=None). Recompute the mean absolute error and compare it to 12.273. It should fall a long way, because the hand-typed numbers were guesses.

Finally, compute the baseline error — predict y.mean() for every row — and check that your fitted model beats it. If it does not, something upstream is wrong.

What to learn next

Researcher — Mathematics and papers.

The pipeline as a composition of stateful transforms

The scikit-learn API encodes a real distinction that hand-written pipelines routinely blur. A transform has two phases:

fit(X_train)        →  estimates parameters θ from training data only
transform(X, θ)     →  a pure function of one row, given θ

Standardisation has θ = (μ, σ) per column. Imputation has θ = the fill value. One-hot encoding has θ = the observed category set. Target encoding has θ = per-category target statistics, and is the most leakage-prone transform in common use.

The invariant that must hold end to end: θ is a function of the training partition only, and transform is applied identically at training and at serving time. Every version of the leakage bug is a violation of one of those two clauses. Expressing the pipeline as a composition, with fitted state carried in one object, makes the invariant structural instead of a thing you remember.

Training-serving skew is the same failure displaced in time: the notebook standardises with pandas, the serving path reimplements it in Java, and the two disagree on ddof or on how an unseen category is handled. Serialising the fitted transform, rather than reimplementing it, is the only robust answer.

Leakage, taxonomised

Kaufman, Rosset and Perlich (2012) give the canonical treatment; Kapoor and Narayanan (2023) document how widespread it remains in published science.

  • Preprocessing leakage. θ estimated over the full dataset before splitting. The example above.
  • Target leakage. A feature that encodes the label, often through a proxy: days_since_claim_settled in a claim-prediction model.
  • Temporal leakage. Random splitting of time-ordered data, letting the model see the future. Requires a forward-chaining split.
  • Group leakage. The same patient, user or document appearing in both partitions. Requires grouped splitting on the entity key.
  • Duplicate leakage. Near-duplicate rows straddling the split. Deduplicate before splitting, not after.

The diagnostic that catches most of these costs nothing: an implausibly high validation score early in a project is evidence of a bug, not of a good model.

Numerical detail that changes results

Estimator bias in scaling. The population variance divides by n, the unbiased sample variance by n − 1. numpy.std defaults to ddof=0; pandas.Series.std defaults to ddof=1; sklearn.preprocessing.StandardScaler uses ddof=0. For n = 8 the two differ by about 7 percent, which is enough to matter for a distance-based model and irrelevant for a tree.

Precision. IEEE 754 float32 gives roughly 7 decimal digits, float64 roughly 16. Deep learning runs in float32 or lower because memory bandwidth, not arithmetic, is the binding constraint; classical tabular ML runs in float64 because it costs nothing at that scale. The hazard sits between them: accumulating a long sum in float32 loses precision quickly, which is why NumPy uses pairwise summation and why mixed-precision training keeps a float32 master copy of the weights while computing in bfloat16.

Feature scale and optimiser conditioning. Standardising is not cosmetic. For gradient descent on a quadratic objective, the convergence rate depends on the condition number of the Hessian, which for a linear model is governed by the spread of the feature covariance eigenvalues. Leaving area in the hundreds beside rooms in single digits inflates that condition number directly. Tree ensembles are invariant to monotone rescaling and need none of this; nearest neighbours, SVMs, ridge and any gradient-trained model need all of it. See optimization.

The cost model, stated plainly

Wall-clock time in a training step divides into interpreter overhead and kernel time. The rule that governs pipeline design:

Python loop over elements   →  pathological; interpreter dispatch dominates
Python loop over batches    →  free; each iteration triggers milliseconds of C

NumPy operations dispatch to compiled ufuncs and, for @, to a BLAS gemm. A matmul of (m, k) @ (k, n) costs 2mkn FLOPs and is compute-bound at large sizes; an elementwise operation over n elements costs O(n) FLOPs against O(n) bytes moved and is memory-bandwidth-bound. That distinction explains why fusing elementwise chains matters and why fusing matmuls does not.

The corollary for pandas: a groupby().agg() is a compiled hash aggregation, while the same logic written with iterrows is a Python loop that also rebuilds a Series per row. The pandas lesson covers the mechanics.

Reproducibility

Four independent sources of nondeterminism, in rough order of how often they bite:

  1. Unseeded RNGs. Prefer numpy.random.Generator (PCG64) created explicitly and threaded through as an argument. The legacy global np.random.seed sets state that any imported library can also mutate.
  2. Hash randomisation. PYTHONHASHSEED varies per process, so set iteration order and anything derived from it varies between runs. Sort before you rely on order.
  3. Floating-point non-associativity. Multi-threaded reductions sum in a nondeterministic order, so results differ in the last bits. OMP_NUM_THREADS=1 makes BLAS reductions deterministic at a real cost in speed.
  4. Library nondeterminism on accelerators. Atomics in GPU kernels; cuDNN algorithm autotuning. torch.use_deterministic_algorithms(True) opts out where a deterministic kernel exists.

Pinning versions matters as much as seeding. A default that changed between minor releases — pandas on fillna downcasting, sklearn on OneHotEncoder(sparse_output=...) — changes results with no code change on your side.

Validation as code, not as vigilance

The failure mode of a data pipeline is a silent wrong answer, so assertions belong inside the pipeline rather than in your memory of it. The cheap set, worth adding to every project:

  • Row count before and after every merge and filter.
  • Expected dtype per column, checked after load.
  • Null fraction per column, with a threshold that fails the run.
  • Range and domain checks on the target, and on any feature with a physical meaning.
  • Train and test schema equality, checked as an assertion, not by eye.

pandera, great_expectations and a handful of assert statements all work. The choice matters much less than having any of them.

References

  • Harris, C. R., et al. "Array programming with NumPy." Nature 585, 357–362, 2020.
  • McKinney, W. "Data Structures for Statistical Computing in Python." Proc. 9th Python in Science Conf., 2010.
  • Pedregosa, F., et al. "Scikit-learn: Machine Learning in Python." JMLR 12, 2825–2830, 2011.
  • Buitinck, L., et al. "API design for machine learning software: experiences from the scikit-learn project." ECML PKDD Workshop, 2013 — the fit/transform contract, argued from first principles.
  • Kaufman, S., Rosset, S., Perlich, C. "Leakage in Data Mining: Formulation, Detection, and Avoidance." ACM TKDD 6(4), 2012.
  • Kapoor, S., Narayanan, A. "Leakage and the reproducibility crisis in machine-learning-based science." Patterns 4(9), 2023.
  • Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." NeurIPS, 2015.

What to learn next

What to learn next

These follow on from what you just read.

  • Mathematics for AI

    Why AI needs mathematics

    You do not need to be a mathematician to build AI. You need four small ideas, mainly so you can tell why a model failed instead of guessing.

  • Mathematics for AI

    Linear algebra

    Linear algebra is the maths of scaling things and adding them up. Every layer of every AI model is that one move, repeated at enormous scale.

  • Mathematics for AI

    Vectors and matrices

    A vector is an ordered list of measurements about one thing. A matrix is many of those lists stacked into a table. Almost everything inside an AI model is one of these two.