Python for AI

Variables and data types

A variable is a name you stick onto a value so you can use it again later. A type is what kind of value it is, and it decides what the computer will let you do with it.

Read these first

On this page 9
  1. Why you should care
  2. Why it exists
  3. The part that catches everyone
  4. The five kinds you need on day one
  5. How it works
  6. A real example you have seen
  7. An honest word
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A variable is a name you stick onto a value, so you can use that value again later.

Think of the jars on a kitchen shelf. Someone sticks a strip of tape on each one and writes on it: sugar, salt, tea. The tape is not the sugar. It is a name for the jar, so anyone can find it without opening every lid.

A variable is that strip of tape. You write a name, you point it at a value, and from then on the name works instead of the value.

You can also peel the tape off one jar and stick it on another. Names in Python move like that too.

Why you should care

Every line of AI code you will ever read is full of names. model, data, learning_rate, epochs, loss.

If names feel comfortable, that code looks like an outline. If they do not, it looks like noise. This is the shortest lesson on the site with the biggest payoff.

Why it exists

Two reasons, and both are practical.

The first is reuse. If you write a value once and name it, you can use it in twenty places. When it needs to change, you change one line instead of hunting for twenty.

The second is memory — yours, not the computer's. total_marks tells the next reader what the value means. The bare number tells them nothing.

The part that catches everyone

Values come in kinds. The kind is called a type, and it decides what the computer will let you do with the value.

Here is the trap, and it is worth reading twice.

Picture a photograph of a five-hundred-rupee note. It looks like money. It is the right colour, the right size, the right number printed on it. You cannot buy chai with it.

The computer sees the same difference. Text that reads "500" is a picture of a number. A number is a number. They look identical on your screen and behave nothing alike.

The five kinds you need on day one

  • Text, called a string in Python. Names, cities, messages. Written inside quotes.
  • Whole numbers, called an int. Counts of things. No decimal point.
  • Numbers with a decimal point, called a float. Prices, weights, scores.
  • True or false, called a bool. Answers to yes-or-no questions.
  • Nothing at all, called None. A box that is deliberately empty.

That is the whole list. There are more, and you will meet them later, and none of them are urgent.

How it works

A name points at a value. That is the entire picture.

   the name              what it points at        its type
   --------              -----------------        --------
   name         →        Asha                     text
   age          →        17                       whole number
   height       →        1.62                     decimal number
   is_student   →        true                     true or false
   scholarship  →        nothing yet              None

Pointing a name somewhere new does not disturb the old value. It moves the tape.

   before a birthday:    age  →  17

   after a birthday:     age  →  18

The number seventeen was not edited. The name stopped pointing at it and started pointing at eighteen.

A real example you have seen

You have filled an online form that refused your phone number because you typed a space in it. Or a shopping site that showed the total as two prices stuck together instead of added up.

Both are type problems. The form wanted a whole number and got text. The shopping site had prices stored as text. It glued them end to end. That is what a computer does with two pieces of text.

Nobody was careless. This is the most common bug in a beginner's first week. It is common because the screen shows you no difference at all.

An honest word

You will hit this bug. Everyone does, including people who have written Python for ten years and are loading a new file at midnight.

The fix is a habit, not a talent. When a number behaves strangely, ask the computer what type it is before you ask anything else. There is a one-word way to do that, and it is in the Developer tab.

Remember this

  • A variable is a name pointed at a value, and you can re-point it whenever you like.
  • A type is what kind of value it is, and it decides what operations are allowed.
  • Text that looks like a number is not a number, and telling them apart early saves hours.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install past Python itself, and nothing to import. Everything here is core Python.

You need somewhere to type it. There are two places, and both are fine.

The interactive prompt. Open Command Prompt on Windows, or Terminal on macOS and Linux. A terminal is a plain window where you type commands as text instead of clicking buttons. Type python there and press Enter. Three arrows appear, >>>. Whatever you type runs the instant you press Enter. Type exit() to leave.

If python is not found, try python3. On Windows, reinstall from python.org and tick "Add Python to PATH".

A saved file. A file is your instructions stored on disk under a name, so they survive after the window closes. Save the code below as profile.py in any folder. In the terminal, move into that folder with cd foldername, then run python profile.py.

Use the prompt for one-liners. Move to a file the moment you have more than three lines to keep.

Your first variables

profile.py
# A profile card. Core Python only - nothing to install, nothing to import.

name = "Asha"          # text, called a string
age = 17               # a whole number, called an int
height_m = 1.62        # a number with a decimal point, called a float
is_student = True      # true or false, called a bool
scholarship = None     # "nothing here yet", a value of its own

print(name, "is", age, "years old")

print("name       ->", type(name).__name__)
print("age        ->", type(age).__name__)
print("height_m   ->", type(height_m).__name__)
print("is_student ->", type(is_student).__name__)
print("scholarship->", type(scholarship).__name__)

age = age + 1          # the same name, re-pointed at a new value
print("after a birthday:", age)
Output
Asha is 17 years old
name       -> str
age        -> int
height_m   -> float
is_student -> bool
scholarship-> NoneType
after a birthday: 18

Line-by-line walkthrough

name = "Asha" binds the name on the left to the value on the right. The single = is not a claim that two things are equal. It is an instruction: point this name at that value.

Quotes decide everything. "17" is text. 17 is a number. Python will not warn you when you pick the wrong one, because both are legal.

type(x) hands back the type object itself, which prints as <class 'int'>. Adding .__name__ pulls out the short label. This is your first debugging tool, and you will reach for it constantly.

None has its own type, printed as NoneType. It means "deliberately empty", which is different from zero and different from an empty piece of text. Missing values in real datasets land here.

age = age + 1 reads oddly the first time. Python works out the right side first, gets 18, then points age at it. The old 17 is discarded.

b = a copies where the name points, not the value. With numbers and text that distinction never bites, because those values cannot be changed in place. With lists it bites hard, and lists and dictionaries shows exactly how.

Converting between types

convert.py
typed = "17"                    # this is what input() always hands you: text

print("as text  :", typed + typed)
print("as number:", int(typed) + int(typed))
print("back to text:", str(int(typed) + int(typed)) + " marks")

price = "12.5"
print("float:", float(price))
print("int of a float:", int(float(price)))   # the .5 is dropped, not rounded

print("bool of an empty string:", bool(""))
print("bool of any other string:", bool("no"))
Output
as text  : 1717
as number: 34
back to text: 34 marks
float: 12.5
int of a float: 12
bool of an empty string: False
bool of any other string: True

Three things in that output deserve a second look.

"17" + "17" gives 1717. The + sign means "add" for numbers and "stick together" for text. Same symbol, two meanings, decided entirely by type.

int(12.5) gives 12. It cuts toward zero and throws the remainder away. int(-12.5) gives -12, not -13. Use round() when you want rounding.

bool("no") is True. Any non-empty piece of text counts as true, including the text "False". Reading a config file and testing the string directly is a classic silent bug.

Decimals are not exact

money.py
print(0.1 + 0.2)
print(0.1 + 0.2 == 0.3)
print(round(0.1 + 0.2, 2) == 0.3)
Output
0.30000000000000004
False
True

This is not a Python flaw. Floats are stored in binary, and one tenth has no exact binary form, in the same way one third has no exact decimal form. Every mainstream language behaves this way.

The practical rule: never test two decimal numbers with ==. Round first, or use math.isclose(a, b). For money, use whole paise as an int, or the decimal module.

Common mistakes

1. Naming a variable after something Python already provides.

python
sum = 10 + 5
print(sum)
print(sum([1, 2, 3]))
Output
15
Traceback (most recent call last):
  File "shadow.py", line 3, in <module>
    print(sum([1, 2, 3]))
TypeError: 'int' object is not callable

Your name replaced the built-in sum function for the rest of the program. The same trap waits behind list, dict, str, type, id and input. Fix: add a word — total, word_list, raw_str.

2. Converting text that is not a whole number.

python
print(int("17.5"))
Output
Traceback (most recent call last):
  File "convert.py", line 1, in <module>
    print(int("17.5"))
ValueError: invalid literal for int() with base 10: '17.5'

int() refuses to guess. Fix: int(float("17.5")) if dropping the decimals is what you want. This error appears the first time you load a real data file, because someone typed 17.5 in a column of whole numbers.

3. A typo in a name.

python
total = 90
print(totl)
Output
Traceback (most recent call last):
  File "typo.py", line 2, in <module>
    print(totl)
NameError: name 'totl' is not defined. Did you mean: 'total'?

Python has no idea you meant total, because you never created totl. Python 3.10 and newer add the helpful Did you mean: 'total'? you see above; older versions stop at the message itself.

A note on reading errors at all. Every traceback on this site is shown in its shortest form. On Python 3.11 and newer you will also see a line of ^ characters under the exact part of the line that failed. That marker is a gift — it points straight at the culprit inside a long line. And always read a traceback from the bottom up: the last line names the problem, and the lines above it show the path that got there.

4. Assuming a value has the type you expect.

Data read from a file, a form, or a web response is text until you convert it. Before debugging a strange result, print type(x) on the value. This one habit resolves a large share of first-month bugs.

Try it yourself

Make three variables for a train ticket: passenger as text, seats as a whole number, and fare as a decimal number. Print a sentence using all three.

Then set seats = "2" with quotes and try to compute the total fare. Read the error, print the type of each value, and fix it with int().

Last, try print(0.1 + 0.2 + 0.3 == 0.6) and predict the answer before you run it.

What to learn next

Researcher — Mathematics and papers.

Names bind to objects

Python has no variables in the C sense. There is no named slot in memory holding a value. There are objects, which live on the heap and carry their own type, and there are names, which are entries in a namespace dictionary that reference objects.

Assignment therefore never copies. b = a adds a second reference to one object. The consequences are exact:

  • Rebinding a name (a = 6) affects only that name.
  • Mutating an object through a name (a.append(1)) is visible through every name bound to it.
  • Immutable types — int, float, str, bytes, tuple, frozenset, bool, NoneType — have no in-place mutation, which is why the distinction stays invisible until you meet a list.

CPython manages object lifetime with reference counting plus a cycle collector for reference cycles. A PyObject header on a 64-bit build is 16 bytes before any payload: an 8-byte refcount and an 8-byte type pointer.

Identity, caching and interning

objects.py
a = 1000
b = int("1000")            # built at run time, so the compiler cannot share it
print("equal values :", a == b)
print("same object  :", a is b)

s = 256
t = int("256")
print("256 is cached:", s is t)   # CPython pre-builds every int from -5 to 256
u = 257
v = int("257")
print("257 is not   :", u is v)

big = 2 ** 100
print("int stays exact  :", big + 1 - big)
print("float rounds off :", float(big) + 1.0 == float(big))
print("float has 53 bits:", 2.0 ** 53 == 2.0 ** 53 + 1.0)
Output
equal values : True
same object  : False
256 is cached: True
257 is not   : False
int stays exact  : 1
float rounds off : True
float has 53 bits: True

Every line above is a CPython implementation detail, not a language guarantee. PyPy, GraalPy and MicroPython are free to differ, and the small-integer range has changed across CPython releases.

Two further sharing mechanisms exist. Equal constants inside a single code object are folded into one entry in co_consts, so a = 1000; b = 1000 written in one module can make a is b true. String literals that look like identifiers are interned automatically; sys.intern forces it for computed strings, which is a real optimisation for dictionary keys parsed from a file.

The operational rule is unchanged: compare values with ==, and reserve is for None, True and False.

Numeric semantics

int is arbitrary precision. CPython stores the magnitude as an array of 30-bit digits with a separate sign. Addition is O(d) in the digit count; multiplication switches from schoolbook O(d²) to Karatsuba above a threshold. Practical consequence: integer arithmetic never overflows, and it silently gets slower as values grow. NumPy's int64 does the opposite — fixed cost, and it wraps around on overflow.

float is IEEE 754 binary64. One sign bit, 11 exponent bits, 52 stored significand bits for 53 bits of precision. Machine epsilon is 2⁻⁵², roughly 2.22e-16. Consecutive representable values near 1.0 differ by about 2.2e-16; near 10⁶ they differ by about 1.2e-10. That drift is why summing a long array left to right accumulates error, and why math.fsum and pairwise summation exist. NumPy uses pairwise summation inside np.sum, so it is measurably more accurate than a Python loop on the same data.

bool subclasses int. True == 1 and isinstance(True, int) are both true. sum([True, False, True]) returns 2, which is exactly how boolean masks get counted in NumPy.

The numeric tower is defined in PEP 3141 via numbers.Number and friends. decimal.Decimal gives configurable-precision base-10 arithmetic with correct rounding modes; fractions.Fraction gives exact rationals. Both are far slower than float and both are correct choices for money.

Memory, and why arrays exist

A Python float object on a 64-bit build costs 24 bytes: the 16-byte object header plus an 8-byte double. A list of a million floats also stores a million 8-byte pointers, and the objects themselves are scattered across the heap.

A NumPy float64 array stores a million raw doubles in one contiguous buffer, with a single type descriptor for the whole array. That is roughly a quarter of the memory, and — more importantly — it is cache-friendly and can be handed straight to a BLAS kernel.

Exact byte counts from sys.getsizeof shift between CPython releases, so measure on your own interpreter rather than trusting a number from a blog. The ratio is the durable fact. This is the concrete reason NumPy exists, and the reason a dtype is a first-class concept there while a type is an afterthought here.

Typing

Python is dynamically typed — types belong to objects, checked at run time — and strongly typed — it will not coerce "1" + 1 into anything.

PEP 484 added annotations. They are stored in __annotations__ and are not enforced at run time; a wrong annotation costs nothing until a checker runs. PEP 563 and PEP 649 changed how annotations are evaluated, which matters mainly for forward references.

In ML codebases annotations pay for themselves on tensor-shaped code, where the interesting invariants are shapes and dtypes rather than Python types. jaxtyping and torchtyping express those; a plain np.ndarray annotation says almost nothing.

References

  • van Rossum, G. Python Language Reference, section 3, "Data model" — the definition of objects, identity and value.
  • Goldberg, D. "What Every Computer Scientist Should Know About Floating-Point Arithmetic." ACM Computing Surveys 23(1), 1991.
  • IEEE 754-2019, Standard for Floating-Point Arithmetic.
  • PEP 3141 — A Type Hierarchy for Numbers, Yasskin, 2007.
  • PEP 484 — Type Hints, van Rossum, Lehtosalo and Langa, 2014.
  • CPython source, Objects/longobject.c and Objects/floatobject.c.

What to learn next

What to learn next

These follow on from what you just read.

  • Python for AI

    Lists and dictionaries

    A list holds many values in order, found by position. A dictionary holds many values by label, found by name. Almost every dataset you will ever load is built from these two.

  • Python for AI

    Loops and conditions

    A loop does the same work once for every item in a group. A condition decides which work to do. Together they are how every program makes a decision about every row of your data.

  • Python for AI

    Functions

    A function is a block of work you give a name, so you can run it again with different inputs. Every AI library you will ever use is a pile of functions someone else already wrote.