Python for AI

Python basics

Python is a way of writing step-by-step instructions for a computer using words close to English. It is the language almost all AI is built in.

On this page 7
  1. Why you should care
  2. Why Python exists
  3. How it works
  4. A real example you have seen
  5. An honest word about the start
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Python is a way of writing step-by-step instructions that a computer will follow, using words close to English.

Think about writing a recipe for someone who has never cooked. You cannot write "make chai" and walk away. You have to write: boil the water, add tea leaves, wait, add milk, strain. Python is that recipe, written for a computer instead of a person.

The computer is a very fast cook with no common sense. It will follow your recipe exactly. If you forget the milk, it serves black tea and never asks why.

Why you should care

Nearly every AI tool you have used was built with Python. ChatGPT, Google Translate, the face unlock on your phone, the recommendations on YouTube. The people who built those things wrote Python.

You do not need to become a great programmer to work with AI. You need enough Python to load some data, call a library, and read the answer. That is a smaller goal than it sounds.

Why Python exists

Early programming languages were written for the machine's comfort, not yours. A working program looked like a wall of symbols and abbreviations. Beginners gave up before they built anything.

In 1991, a Dutch programmer named Guido van Rossum released Python with a different priority. Code is read far more often than it is written, so reading it should feel pleasant. He removed the punctuation clutter and used blank space to show structure.

That decision is why a Python program looks close to an outline you would write on paper.

How it works

You write a file      Python reads it        The computer
of instructions   →   one line at a time  →  does what each   →   You see
called script.py      from top to bottom     line asks               the result

There is no in-between step. You save the file, you run it, you see the answer. That fast loop is why Python is a comfortable place to learn.

Three ideas carry most of the language:

STORE    →  give a value a name, so you can use it later
DECIDE   →  do this thing only when some condition is true
REPEAT   →  do this thing once for every item in a group

Everything else in Python is a variation on those three.

A real example you have seen

When you upload a photo to Google Photos and later search "beach", it finds the beach pictures. Nobody labelled your photos by hand. A model looked at each one and worked out what was in it.

The people who trained that model wrote their instructions in Python. They loaded thousands of pictures, described the model shape, and told it to learn. Their file looked far more like an outline than like a wall of symbols.

An honest word about the start

The hardest part of week one is usually not the language. It is getting Python installed and getting your first file to run. Many people quit at that step and think they are bad at coding.

They are not. Setup is genuinely fiddly, and it is fiddly for professionals too. Push through that one afternoon, and the rest gets much friendlier.

Remember this

  • Python is a set of written instructions, followed exactly and in order.
  • It was designed to be readable by humans first, which is why beginners can start with it.
  • You need only a small slice of Python to begin doing real AI work.

What to learn next

Developer — Code and libraries.

Setup

Core Python needs no package installed. Check what you have:

bash
python --version

If that prints Python 3.10 or higher, you are ready. If the command is not found, install from python.org and tick "Add Python to PATH" on Windows.

Everything in this lesson runs on the standard library. No pip install required.

Minimal runnable code

Save this as grades.py and run python grades.py.

grades.py
# A tiny grade book. Everything here is core Python - nothing to install.

students = ["Asha", "Ravi", "Meera"]
scores = [88, 71, 95]

def grade(score):
    # Early returns keep the common case at the top and avoid deep nesting.
    if score >= 90:
        return "A"
    if score >= 80:
        return "B"
    return "C"

total = 0
for name, score in zip(students, scores):
    total = total + score
    print(f"{name:<6}{score:>4}  grade {grade(score)}")

average = total / len(students)
print(f"Class average: {average:.1f}")
print("Topper:", students[scores.index(max(scores))])
Output
Asha    88  grade B
Ravi    71  grade C
Meera   95  grade A
Class average: 84.7
Topper: Meera

Line-by-line walkthrough

students = [...] creates a list, which is an ordered collection you can grow and change. Square brackets mean list.

def grade(score): defines a function, a named block of code you can call many times. Everything indented under it belongs to it. The indentation is not decoration — it is the syntax.

zip(students, scores) walks two lists side by side and hands you one pair at a time. Without it you would index into both lists by position, which is a common source of bugs.

f"{name:<6}{score:>4}" is an f-string, a string that can contain values inside braces. The <6 pads the name to six characters aligned left. The >4 pads the score to four characters aligned right. That is what produces the neat columns above.

{average:.1f} rounds the display to one decimal place. It changes the printed text, not the stored value.

scores.index(max(scores)) finds the highest score, then finds its position, then uses that position to pick the matching name. Read the inner call first.

Common mistakes

1. Mixing tabs and spaces for indentation. Python counts leading whitespace, and a tab is not four spaces to the interpreter.

Output
TabError: inconsistent use of tabs and spaces in indentation

Fix: set your editor to insert 4 spaces when you press Tab. Every serious editor has this setting.

2. Using = where you meant ==. One equals sign assigns a value. Two compare values.

python
if score = 90:      # SyntaxError
if score == 90:     # correct

3. Forgetting that input() always gives you text.

python
age = input("age: ")     # you type 17, but age holds the text "17"
print(age + 1)
Output
age: 17
Traceback (most recent call last):
  File "age.py", line 2, in <module>
    print(age + 1)
TypeError: can only concatenate str (not "int") to str

Fix: convert it with age = int(input("age: ")). This bites everyone once, because the text "17" and the number 17 look identical when printed.

4. Integer division surprises. 7 / 2 gives 3.5, but 7 // 2 gives 3. The second one throws away the remainder. Data bugs from // are quiet and painful, because nothing crashes.

Try it yourself

Add a fourth student with a score of 90. Then change grade() so a score of exactly 90 returns "A+" instead of "A". Run it and confirm only the new student changes.

Then try removing one name from students without removing its score. Notice that zip() stops at the shorter list and says nothing. Silence like that is why professionals check their data lengths.

What to learn next

  • Lists and dictionaries — the two containers you will use daily.
  • Functions — arguments, defaults, and return values in depth.
  • NumPy — where Python starts being fast enough for AI.

Researcher — Mathematics and papers.

What actually runs

CPython, the reference implementation, does not execute your source text. It compiles each module to bytecode, a compact instruction set for a stack machine. A C loop in ceval.c then interprets those instructions.

You can read the bytecode with the standard library:

peek.py
import dis

def add_tax(price):
    return price * 1.18

dis.dis(add_tax)
Output
  4           0 LOAD_FAST                0 (price)
              2 LOAD_CONST               1 (1.18)
              4 BINARY_MULTIPLY
              6 RETURN_VALUE

That output is from CPython 3.10. The opcode set is an internal detail and changes between releases. From 3.11 onward you will see extra opcodes such as RESUME, and BINARY_MULTIPLY was folded into BINARY_OP. Read the shape, not the exact listing.

The important observation: four bytecode instructions to multiply two numbers. Each one costs a dispatch through the interpreter loop, plus reference-count updates on boxed objects. A C compiler would emit roughly one machine instruction. That gap is the whole story of Python performance.

Cost of common operations

Amortised time complexity for the built-in types, where n is the number of elements:

Operationlistdict / set
index or key lookupO(1)O(1) average
append / insert at endO(1) amortisedO(1) average
insert or delete at frontO(n)not applicable
x in containerO(n)O(1) average
iterate allO(n)O(n)

The x in list row is the classic accidental quadratic. A loop over n items that tests membership against a list of n items is O(n²). Changing the inner container to a set makes it O(n). This shows up constantly in text-preprocessing code.

list.append is amortised O(1) because CPython over-allocates the backing array. The growth factor is roughly 1.125 plus a constant, so reallocation is rare.

The GIL

CPython protects its object memory with a global interpreter lock. That is a single mutex letting only one thread execute bytecode at a time. Threads still help when the work is waiting on disk or network, because the lock is released during those calls. Threads do not help pure Python computation.

The standard escapes are:

  • multiprocessing, which pays a serialisation cost per message but gets real parallelism.
  • Releasing the GIL inside C extensions. NumPy, PyTorch, and BLAS kernels do exactly this, which is why a matrix multiply already uses every core.
  • Free-threaded builds. PEP 703 (Sam Gross, 2023) proposed making the GIL optional. CPython 3.13 shipped an experimental free-threaded build in 2024. Treat it as experimental rather than production-ready.

Why AI uses a slow language

The apparent contradiction resolves once you measure where the time goes. In a typical training step, well under one percent of wall-clock time is spent in the CPython interpreter. The rest is inside compiled kernels. BLAS routines called by NumPy. cuBLAS and cuDNN called by PyTorch. Fused CUDA kernels emitted by a compiler.

Python is the orchestration layer, meaning it decides what happens and in what order, while compiled code does the arithmetic. The cost model is therefore:

  • Python-level loop over array elements — pathological, avoid.
  • Python-level loop over batches — negligible, because each iteration triggers milliseconds of C work.

This is also why torch.compile and CUDA graphs exist. Sometimes per-step Python overhead grows comparable to kernel time. That happens with very small models, or with very large graphs. The fix is to capture the launches ahead of time.

Semantics worth being precise about

Names bind to objects; they do not label memory slots. a = b copies a reference. Mutating through a is visible through b. Every "why did my list change" bug is this rule.

Default arguments are evaluated once, at function definition time. def f(x, seen=[]) shares one list across all calls. Use seen=None and build inside the body.

Small integers and short strings are interned. 256 is 256 is True while 257 is 257 may be False depending on how the constants were compiled. Compare values with ==, and reserve is for None.

Everything is late-bound. Attribute lookup consults the instance dictionary, then the class, then the MRO, at every access. This is what makes monkey-patching possible and micro-optimisation hard.

References

  • van Rossum, G. Python Reference Manual, CWI, 1995. The original language definition.
  • van Rossum, G., Warsaw, B., Coghlan, N. PEP 8 — Style Guide for Python Code, 2001, revised continuously.
  • Gross, S. PEP 703 — Making the Global Interpreter Lock Optional in CPython, 2023.
  • CPython source, Python/ceval.c and Objects/listobject.c. The list growth factor is in list_resize.
  • Harris, C. R., et al. "Array programming with NumPy." Nature 585, 357–362, 2020. Background for why the compute leaves Python.

What to learn next

  • NumPy — the array model that moves computation out of the interpreter.
  • Functions — closures, defaults, and argument binding.
  • Python for machine learning — the subset of the language that ML code actually uses.

What to learn next

These follow on from what you just read.

  • Python for AI

    Variables and data types

    A variable is a name you stick onto a value so you can use it again later. A type is what kind of value it is, and it decides what the computer will let you do with it.

  • Python for AI

    Lists and dictionaries

    A list holds many values in order, found by position. A dictionary holds many values by label, found by name. Almost every dataset you will ever load is built from these two.

  • Python for AI

    Loops and conditions

    A loop does the same work once for every item in a group. A condition decides which work to do. Together they are how every program makes a decision about every row of your data.