Python for AI

Functions

A function is a block of work you give a name, so you can run it again with different inputs. Every AI library you will ever use is a pile of functions someone else already wrote.

On this page 9
  1. Why you should care
  2. Why they exist
  3. How it works
  4. The one thing beginners get wrong
  5. Default answers
  6. A real example you have seen
  7. An honest word
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A function is a block of work with a name on it, so you can run that work again whenever you like.

Think of the mixer-grinder in a kitchen. You drop coconut, chilli and salt into the jar, press the button, and chutney comes out. You do not rebuild the machine each time.

Tomorrow you put tomato and garlic in the same jar and press the same button. Different things in, different thing out, identical machine.

That is a function. What you put in are the inputs. What comes out is the result. The steps in between were written once.

Why you should care

Every line of AI code you will read is calls to functions. train(), predict(), read_csv(), fit().

You do not need to write clever ones. You need to be comfortable reading them, passing things into them, and catching what comes out. That is a small skill with an enormous return.

Why they exist

You have already seen the alternative, and it hurts in three ways.

Repetition. The same eight lines appear in four places. Fixing a bug means fixing it four times, and you will forget one.

No name. Eight lines of arithmetic sitting in the middle of a file explain nothing. average_marks(scores) explains itself.

No boundary. Without functions, every name in your program can touch every other name. A function draws a line: what goes in, what comes out, nothing else.

How it works

   what goes in                 the machine                 what comes out
   ------------                 -----------                 --------------

   coconut, chilli, salt   →    grind for thirty       →    chutney
                                seconds

   tomato, garlic          →    the same machine       →    tomato chutney
                                the same button

You write the machine once. You use it as many times as you like, with different things going in each time.

The one thing beginners get wrong

A function can do two different things with its answer, and they are not the same.

Picture the juice stall again. The man can hand you the glass, or he can shout the price across the road. Both feel like an answer. Only one of them can be drunk.

   printing  →  shows the answer on the screen, and hands back nothing
   returning →  hands the answer back, so the rest of your program can use it

If you print when you meant to return, the screen looks right and the program does nothing useful with the value. This confuses almost everyone at the start, and once it clicks it never confuses you again.

Default answers

Some inputs have a sensible normal value. A pass mark is usually forty.

You can give an input a default, so anyone using the function can leave it out. They get the normal behaviour with less typing. They can still override it when a particular exam sets a different pass mark.

A real example you have seen

Every button in every app is a function. Tap "pay", and one named block of work runs: check the balance, deduct, send a message, update the screen.

Nobody rewrote those steps for each user. It was written once and it runs a few million times a day.

When you later write model.fit(data), you are pressing a button someone else built. Inside are thousands of lines you will never need to read.

An honest word

Naming things is the hard part, not the syntax. process() and do_stuff() tell a future reader nothing, and the future reader is usually you, three weeks later.

A good name says what comes out. average_marks, clean_text, is_spam. If you cannot name a function, that is usually a sign it is doing two jobs and wants to be two functions.

Remember this

  • A function is named work you can run again, with different inputs each time.
  • Returning hands the answer back; printing only shows it. Learn the difference early.
  • An input can have a default, so common cases stay short.

What to learn next

Developer — Code and libraries.

Setup

Core Python. Nothing to install, nothing to import.

Two words carry the syntax. def starts a definition, and return hands a value back. Everything indented under def belongs to the function.

Three small functions

marks.py
def average(numbers):
    """Return the mean of a list of numbers."""      # a docstring: what it does
    return sum(numbers) / len(numbers)


def band(score, pass_mark=40):        # pass_mark has a default, so callers may omit it
    if score >= 90:
        return "distinction"
    if score >= pass_mark:            # reaching here means the first return did not run
        return "pass"
    return "needs help"


def report(name, scores):
    avg = average(scores)             # functions call other functions
    return f"{name:<6} avg {avg:5.1f}  {band(avg)}"


print(report("Asha", [88, 91, 95]))
print(report("Ravi", [45, 38, 52]))
print("with a stricter pass mark:", band(45, pass_mark=50))
print("average alone:", average([10, 20, 30]))
Output
Asha   avg  91.3  distinction
Ravi   avg  45.0  pass
with a stricter pass mark: needs help
average alone: 20.0

Line-by-line walkthrough

def average(numbers): creates the function but runs nothing. The body executes only when something calls average(...). Defining is not running.

numbers is a parameter — a name that exists only inside the function. The value handed in at the call site is the argument. Two words for the two halves of the same connection.

The docstring is a string on the first line of the body. help(average) prints it, and every editor shows it when you hover the name. Treat it as the smallest possible piece of documentation and write it.

return ends the function immediately. The three returns in band are not an oversight. Once one fires, the rest of the body never runs, which is why band needs no elif. This early-return style keeps the common case at the top and avoids deep nesting.

pass_mark=40 is a default argument. Calling band(45) uses forty. Calling band(45, pass_mark=50) overrides it. Naming the argument at the call site costs six characters and makes the line readable a year later.

band(avg) inside report shows the useful part: small functions compose. Each one is testable on its own and readable on its own.

Return, print, and scope

return_vs_print.py
def shout_print(text):
    print(text.upper())          # shows something, hands back nothing


def shout_return(text):
    return text.upper()          # hands the value back to whoever called


a = shout_print("hello")
b = shout_return("hello")

print("what shout_print gave back:", a)
print("what shout_return gave back:", b)
print("only a returned value can be reused:", b + "!!!")

count = 0                        # a name outside every function


def add_one():
    count = 1                    # a NEW name, local to this function only
    return count


print("inside said:", add_one(), "| outside is still:", count)
Output
HELLO
what shout_print gave back: None
what shout_return gave back: HELLO
only a returned value can be reused: HELLO!!!
inside said: 1 | outside is still: 0

A function with no return hands back None. That is not an error and Python will not warn you. It becomes an error later, in a different line, when you try to use the None.

The last two lines show scope. Names created inside a function are born when it is called and vanish when it returns. The count inside add_one never touched the count outside. That isolation is the point of a function, not a limitation of it.

Common mistakes

1. Computing the answer and forgetting to return it.

python
def add(a, b):
    total = a + b      # computed, then thrown away when the function ends


result = add(2, 3)
print(result)
print(result * 2)
Output
None
Traceback (most recent call last):
  File "add.py", line 7, in <module>
    print(result * 2)
TypeError: unsupported operand type(s) for *: 'NoneType' and 'int'

Note where the crash happened: line 7, not line 2. A missing return reports itself far away from the bug. When you see NoneType in an error, look for a function that forgot to return.

2. A list or dictionary as a default value.

python
def add_item(item, basket=[]):
    basket.append(item)
    return basket


print(add_item("idli"))
print(add_item("vada"))
Output
['idli']
['idli', 'vada']

The second call was meant to start empty. Default values are built once, when the def line runs, not once per call. So every call shares one list.

Fix: use None as the marker and build inside the body.

python
def add_item(item, basket=None):
    if basket is None:
        basket = []
    basket.append(item)
    return basket

3. Assigning to an outer name inside a function.

python
total = 0


def bump():
    total = total + 1


bump()
Output
Traceback (most recent call last):
  File "scope.py", line 8, in <module>
    bump()
  File "scope.py", line 5, in bump
    total = total + 1
UnboundLocalError: cannot access local variable 'total' where it is not associated with a value

Because total is assigned somewhere in the body, Python treats it as local for the whole function — including the read on the right-hand side, which happens before anything is stored. On Python 3.10 and older the message reads local variable 'total' referenced before assignment instead.

Fix: return the new value and let the caller rebind it. global also works and is almost always the wrong choice, because it makes the function impossible to reason about alone.

4. Changing an argument in place.

python
def drop_first(rows):
    del rows[0]
    return rows


data = ["a", "b", "c"]
trimmed = drop_first(data)
print("returned:", trimmed)
print("original:", data)
Output
returned: ['b', 'c']
original: ['b', 'c']

The caller's list was modified. The function received a reference, not a copy, exactly as described in variables and data types.

Sometimes that is what you want. When it is not, copy first with rows = rows[:], or name the function so the effect is obvious — drop_first_in_place. Silent mutation of an argument is the bug that costs the most time in data pipelines, because the damage shows up in an unrelated file.

Try it yourself

Write total_bill(qty, price, gst=0.05) that returns the amount payable including tax. Call it once using the default and once with gst=0.18.

Then write clean(text) that strips spaces from both ends and lowercases the result, and use it on " Filter Coffee ". Print the original afterwards and confirm it is unchanged — text cannot be mutated, so this one is safe by construction.

Last, take the pass-counting loop you wrote in loops and conditions and move it into a function called count_passed(scores, pass_mark=40). Notice how the calling code shrinks to one readable line.

What to learn next

Researcher — Mathematics and papers.

Functions are objects

def is an executable statement. It builds a function object at run time and binds it to a name in the enclosing namespace. The object carries __code__ (a compiled code object), __defaults__, __kwdefaults__, __closure__, __globals__, __annotations__ and __dict__.

Because they are ordinary objects, functions can be stored in lists, passed as arguments, returned, and given attributes. That is what makes decorators, callbacks, key= arguments and dependency injection possible without any special language machinery.

Argument binding

The full parameter grammar, in order:

def f(pos_only, /, standard, *args, kw_only, **kwargs)
  • Everything before / is positional-only (PEP 570, Python 3.8). This is how many C builtins have always behaved.
  • Everything after a bare * or after *args is keyword-only (PEP 3102).
  • *args collects surplus positionals into a tuple; **kwargs collects surplus keywords into a dict.
functions.py
from functools import lru_cache


def f(a, b=2, *rest, key=None, **extra):
    return a, b, rest, key, extra


print(f(1))
print(f(1, 3, 4, 5, key="k", flag=True))
print("defaults live on the object:", f.__defaults__, f.__kwdefaults__)

fs = [lambda: i for i in range(3)]         # late binding: i is read at call time
print("late binding :", [g() for g in fs])
fs = [lambda i=i: i for i in range(3)]     # bound at definition time instead
print("early binding:", [g() for g in fs])

calls = 0


@lru_cache(maxsize=None)
def fib(n):
    global calls
    calls += 1
    return n if n < 2 else fib(n - 1) + fib(n - 2)


print("fib(30):", fib(30), "| bodies executed:", calls)
Output
(1, 2, (), None, {})
(1, 3, (4, 5), 'k', {'flag': True})
defaults live on the object: (2,) {'key': None}
late binding : [2, 2, 2]
early binding: [0, 1, 2]
fib(30): 832040 | bodies executed: 31

Three results in that output are worth dwelling on.

Defaults are stored on the function object, evaluated once when def executes. That is the whole explanation of the mutable-default trap, and it is also why def f(x, t=time.time()) freezes a timestamp at import.

Late binding in closures. A closure captures the variable, not the value at capture time, via a cell object. All three lambdas share one cell holding i, which is 2 after the comprehension finishes. The i=i trick works because default values are evaluated eagerly. This is a live bug generator in code that builds a list of per-layer or per-feature callbacks in a loop.

Memoisation collapses exponential recursion. Naive fib(30) executes 2,692,537 bodies; with lru_cache it executes 31. The cache key is the argument tuple, so arguments must be hashable — a real constraint when the argument is a NumPy array or a DataFrame.

Frames, closures and cost

Each Python-level call builds a frame holding the local variable array, the value stack, and a pointer to the caller. Before 3.11 that meant heap-allocating a frame object per call. PEP 659's work, plus the frame changes in 3.11, inlined Python-to-Python calls into the eval loop, removing most of the per-call C stack push and lazily materialising frame objects only when introspection demands one. Calls got substantially cheaper; they did not become free.

The practical rule is unchanged: a Python function called once per array element is a performance defect, and a Python function called once per batch is free. This is the same cost boundary described in loops and conditions.

Scope resolution follows LEGB — local, enclosing, global, builtin. The compiler decides local versus non-local statically, from whether the name is assigned anywhere in the body, which is what produces UnboundLocalError rather than a fall-through to the global. nonlocal (PEP 3104) rebinds in the nearest enclosing function scope; global rebinds at module scope.

Decorators

A decorator is a function taking a function and returning a replacement. @d above def f is exactly f = d(f), evaluated at definition time.

The one obligation is preserving metadata. functools.wraps copies __name__, __doc__, __module__, __qualname__ and __wrapped__ onto the wrapper, without which every decorated function in your codebase reports itself as wrapper in tracebacks and documentation.

Decorators earn their place in ML code for cross-cutting concerns: @lru_cache for expensive pure lookups, @torch.no_grad() for inference, @functools.singledispatch for type-based dispatch, @dataclass for config objects, and timing or retry wrappers around data loading.

Purity, and why it matters here

A pure function depends only on its arguments and mutates nothing outside itself. Purity is not an aesthetic preference in machine-learning code; it is what makes four practical things possible.

  • Reproducibility. A preprocessing step that reads a global or mutates its input cannot be rerun to the same result.
  • Caching. Memoisation is sound only for pure functions.
  • Parallelism. multiprocessing and joblib serialise arguments across process boundaries, so hidden shared state either breaks or silently diverges.
  • Testing. A pure function needs no fixtures.

The two impure things that matter most in practice are global random state and in-place mutation of arrays. Pass an explicit numpy.random.Generator rather than calling np.random.seed, and prefer returning a new array over writing through a view whose owner you cannot see.

Typing

PEP 484 annotations on functions are the highest-value place to use them, because the signature is the contract. PEP 604 allows int | None instead of Optional[int]. typing.Callable, ParamSpec (PEP 612) and TypeVar let a decorator preserve its wrapped signature for a checker.

Annotations are not enforced at run time. pydantic does enforce them, at a real cost, which is why it appears at the edges of a system — parsing configs and API payloads — rather than in inner loops.

References

  • PEP 3102 — Keyword-Only Arguments, Talin, 2006.
  • PEP 3104 — Access to Names in Outer Scopes, Yee, 2006.
  • PEP 318 — Decorators for Functions and Methods, Smith, Coghlan et al., 2003.
  • PEP 570 — Python Positional-Only Parameters, Hastings, Galindo et al., 2018.
  • PEP 612 — Parameter Specification Variables, Lee, 2019.
  • PEP 659 — Specializing Adaptive Interpreter, Shannon, 2021.
  • CPython source, Objects/funcobject.c and Python/ceval.c.

What to learn next

What to learn next

These follow on from what you just read.

  • Python for AI

    NumPy

    NumPy lets you do one operation to millions of numbers at once instead of one at a time. It is the foundation every AI library in Python is built on.

  • Python for AI

    Pandas

    Pandas is a table with named columns that you can filter, group and summarise in one line. It is where almost every AI project starts, because real data arrives as a table.

  • Python for AI

    Matplotlib

    Matplotlib turns your numbers into pictures. Looking at data before modelling it is the single habit that catches the most mistakes.