Functions
A function is a block of work you give a name, so you can run it again with different inputs. Every AI library you will ever use is a pile of functions someone else already wrote.
- 15 min read
- 3 reading levels
- Published
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A function is a block of work with a name on it, so you can run that work again whenever you like.
Think of the mixer-grinder in a kitchen. You drop coconut, chilli and salt into the jar, press the button, and chutney comes out. You do not rebuild the machine each time.
Tomorrow you put tomato and garlic in the same jar and press the same button. Different things in, different thing out, identical machine.
That is a function. What you put in are the inputs. What comes out is the result. The steps in between were written once.
Why you should care
Every line of AI code you will read is calls to functions. train(), predict(), read_csv(), fit().
You do not need to write clever ones. You need to be comfortable reading them, passing things into them, and catching what comes out. That is a small skill with an enormous return.
Why they exist
You have already seen the alternative, and it hurts in three ways.
Repetition. The same eight lines appear in four places. Fixing a bug means fixing it four times, and you will forget one.
No name. Eight lines of arithmetic sitting in the middle of a file explain nothing. average_marks(scores) explains itself.
No boundary. Without functions, every name in your program can touch every other name. A function draws a line: what goes in, what comes out, nothing else.
How it works
what goes in the machine what comes out
------------ ----------- --------------
coconut, chilli, salt → grind for thirty → chutney
seconds
tomato, garlic → the same machine → tomato chutney
the same buttonYou write the machine once. You use it as many times as you like, with different things going in each time.
The one thing beginners get wrong
A function can do two different things with its answer, and they are not the same.
Picture the juice stall again. The man can hand you the glass, or he can shout the price across the road. Both feel like an answer. Only one of them can be drunk.
printing → shows the answer on the screen, and hands back nothing
returning → hands the answer back, so the rest of your program can use itIf you print when you meant to return, the screen looks right and the program does nothing useful with the value. This confuses almost everyone at the start, and once it clicks it never confuses you again.
Default answers
Some inputs have a sensible normal value. A pass mark is usually forty.
You can give an input a default, so anyone using the function can leave it out. They get the normal behaviour with less typing. They can still override it when a particular exam sets a different pass mark.
A real example you have seen
Every button in every app is a function. Tap "pay", and one named block of work runs: check the balance, deduct, send a message, update the screen.
Nobody rewrote those steps for each user. It was written once and it runs a few million times a day.
When you later write model.fit(data), you are pressing a button someone else built. Inside are thousands of lines you will never need to read.
An honest word
Naming things is the hard part, not the syntax. process() and do_stuff() tell a future reader nothing, and the future reader is usually you, three weeks later.
A good name says what comes out. average_marks, clean_text, is_spam. If you cannot name a function, that is usually a sign it is doing two jobs and wants to be two functions.
Remember this
- A function is named work you can run again, with different inputs each time.
- Returning hands the answer back; printing only shows it. Learn the difference early.
- An input can have a default, so common cases stay short.
What to learn next
- NumPy — a library of functions that work on whole piles of numbers.
- Python for machine learning — the shape of a real ML script.
- What is machine learning? — what all these functions are for.
Developer — Code and libraries.
Setup
Core Python. Nothing to install, nothing to import.
Two words carry the syntax. def starts a definition, and return hands a value back. Everything indented under def belongs to the function.
Three small functions
def average(numbers):
"""Return the mean of a list of numbers.""" # a docstring: what it does
return sum(numbers) / len(numbers)
def band(score, pass_mark=40): # pass_mark has a default, so callers may omit it
if score >= 90:
return "distinction"
if score >= pass_mark: # reaching here means the first return did not run
return "pass"
return "needs help"
def report(name, scores):
avg = average(scores) # functions call other functions
return f"{name:<6} avg {avg:5.1f} {band(avg)}"
print(report("Asha", [88, 91, 95]))
print(report("Ravi", [45, 38, 52]))
print("with a stricter pass mark:", band(45, pass_mark=50))
print("average alone:", average([10, 20, 30]))Asha avg 91.3 distinction Ravi avg 45.0 pass with a stricter pass mark: needs help average alone: 20.0
Line-by-line walkthrough
def average(numbers): creates the function but runs nothing. The body executes only when something calls average(...). Defining is not running.
numbers is a parameter — a name that exists only inside the function. The value handed in at the call site is the argument. Two words for the two halves of the same connection.
The docstring is a string on the first line of the body. help(average) prints it, and every editor shows it when you hover the name. Treat it as the smallest possible piece of documentation and write it.
return ends the function immediately. The three returns in band are not an oversight. Once one fires, the rest of the body never runs, which is why band needs no elif. This early-return style keeps the common case at the top and avoids deep nesting.
pass_mark=40 is a default argument. Calling band(45) uses forty. Calling band(45, pass_mark=50) overrides it. Naming the argument at the call site costs six characters and makes the line readable a year later.
band(avg) inside report shows the useful part: small functions compose. Each one is testable on its own and readable on its own.
Return, print, and scope
def shout_print(text):
print(text.upper()) # shows something, hands back nothing
def shout_return(text):
return text.upper() # hands the value back to whoever called
a = shout_print("hello")
b = shout_return("hello")
print("what shout_print gave back:", a)
print("what shout_return gave back:", b)
print("only a returned value can be reused:", b + "!!!")
count = 0 # a name outside every function
def add_one():
count = 1 # a NEW name, local to this function only
return count
print("inside said:", add_one(), "| outside is still:", count)HELLO what shout_print gave back: None what shout_return gave back: HELLO only a returned value can be reused: HELLO!!! inside said: 1 | outside is still: 0
A function with no return hands back None. That is not an error and Python will not warn you. It becomes an error later, in a different line, when you try to use the None.
The last two lines show scope. Names created inside a function are born when it is called and vanish when it returns. The count inside add_one never touched the count outside. That isolation is the point of a function, not a limitation of it.
Common mistakes
1. Computing the answer and forgetting to return it.
def add(a, b):
total = a + b # computed, then thrown away when the function ends
result = add(2, 3)
print(result)
print(result * 2)None
Traceback (most recent call last):
File "add.py", line 7, in <module>
print(result * 2)
TypeError: unsupported operand type(s) for *: 'NoneType' and 'int'Note where the crash happened: line 7, not line 2. A missing return reports itself far away from the bug. When you see NoneType in an error, look for a function that forgot to return.
2. A list or dictionary as a default value.
def add_item(item, basket=[]):
basket.append(item)
return basket
print(add_item("idli"))
print(add_item("vada"))['idli'] ['idli', 'vada']
The second call was meant to start empty. Default values are built once, when the def line runs, not once per call. So every call shares one list.
Fix: use None as the marker and build inside the body.
def add_item(item, basket=None):
if basket is None:
basket = []
basket.append(item)
return basket3. Assigning to an outer name inside a function.
total = 0
def bump():
total = total + 1
bump()Traceback (most recent call last):
File "scope.py", line 8, in <module>
bump()
File "scope.py", line 5, in bump
total = total + 1
UnboundLocalError: cannot access local variable 'total' where it is not associated with a valueBecause total is assigned somewhere in the body, Python treats it as local for the whole function — including the read on the right-hand side, which happens before anything is stored. On Python 3.10 and older the message reads local variable 'total' referenced before assignment instead.
Fix: return the new value and let the caller rebind it. global also works and is almost always the wrong choice, because it makes the function impossible to reason about alone.
4. Changing an argument in place.
def drop_first(rows):
del rows[0]
return rows
data = ["a", "b", "c"]
trimmed = drop_first(data)
print("returned:", trimmed)
print("original:", data)returned: ['b', 'c'] original: ['b', 'c']
The caller's list was modified. The function received a reference, not a copy, exactly as described in variables and data types.
Sometimes that is what you want. When it is not, copy first with rows = rows[:], or name the function so the effect is obvious — drop_first_in_place. Silent mutation of an argument is the bug that costs the most time in data pipelines, because the damage shows up in an unrelated file.
Try it yourself
Write total_bill(qty, price, gst=0.05) that returns the amount payable including tax. Call it once using the default and once with gst=0.18.
Then write clean(text) that strips spaces from both ends and lowercases the result, and use it on " Filter Coffee ". Print the original afterwards and confirm it is unchanged — text cannot be mutated, so this one is safe by construction.
Last, take the pass-counting loop you wrote in loops and conditions and move it into a function called count_passed(scores, pass_mark=40). Notice how the calling code shrinks to one readable line.
What to learn next
- NumPy — functions that operate on entire arrays at once.
- Python for machine learning — assembling these pieces into a pipeline.
- Pandas — passing functions to
.apply()and friends.
Researcher — Mathematics and papers.
Functions are objects
def is an executable statement. It builds a function object at run time and binds it to a name in the enclosing namespace. The object carries __code__ (a compiled code object), __defaults__, __kwdefaults__, __closure__, __globals__, __annotations__ and __dict__.
Because they are ordinary objects, functions can be stored in lists, passed as arguments, returned, and given attributes. That is what makes decorators, callbacks, key= arguments and dependency injection possible without any special language machinery.
Argument binding
The full parameter grammar, in order:
def f(pos_only, /, standard, *args, kw_only, **kwargs)- Everything before
/is positional-only (PEP 570, Python 3.8). This is how many C builtins have always behaved. - Everything after a bare
*or after*argsis keyword-only (PEP 3102). *argscollects surplus positionals into a tuple;**kwargscollects surplus keywords into a dict.
from functools import lru_cache
def f(a, b=2, *rest, key=None, **extra):
return a, b, rest, key, extra
print(f(1))
print(f(1, 3, 4, 5, key="k", flag=True))
print("defaults live on the object:", f.__defaults__, f.__kwdefaults__)
fs = [lambda: i for i in range(3)] # late binding: i is read at call time
print("late binding :", [g() for g in fs])
fs = [lambda i=i: i for i in range(3)] # bound at definition time instead
print("early binding:", [g() for g in fs])
calls = 0
@lru_cache(maxsize=None)
def fib(n):
global calls
calls += 1
return n if n < 2 else fib(n - 1) + fib(n - 2)
print("fib(30):", fib(30), "| bodies executed:", calls)(1, 2, (), None, {})
(1, 3, (4, 5), 'k', {'flag': True})
defaults live on the object: (2,) {'key': None}
late binding : [2, 2, 2]
early binding: [0, 1, 2]
fib(30): 832040 | bodies executed: 31Three results in that output are worth dwelling on.
Defaults are stored on the function object, evaluated once when def executes. That is the whole explanation of the mutable-default trap, and it is also why def f(x, t=time.time()) freezes a timestamp at import.
Late binding in closures. A closure captures the variable, not the value at capture time, via a cell object. All three lambdas share one cell holding i, which is 2 after the comprehension finishes. The i=i trick works because default values are evaluated eagerly. This is a live bug generator in code that builds a list of per-layer or per-feature callbacks in a loop.
Memoisation collapses exponential recursion. Naive fib(30) executes 2,692,537 bodies; with lru_cache it executes 31. The cache key is the argument tuple, so arguments must be hashable — a real constraint when the argument is a NumPy array or a DataFrame.
Frames, closures and cost
Each Python-level call builds a frame holding the local variable array, the value stack, and a pointer to the caller. Before 3.11 that meant heap-allocating a frame object per call. PEP 659's work, plus the frame changes in 3.11, inlined Python-to-Python calls into the eval loop, removing most of the per-call C stack push and lazily materialising frame objects only when introspection demands one. Calls got substantially cheaper; they did not become free.
The practical rule is unchanged: a Python function called once per array element is a performance defect, and a Python function called once per batch is free. This is the same cost boundary described in loops and conditions.
Scope resolution follows LEGB — local, enclosing, global, builtin. The compiler decides local versus non-local statically, from whether the name is assigned anywhere in the body, which is what produces UnboundLocalError rather than a fall-through to the global. nonlocal (PEP 3104) rebinds in the nearest enclosing function scope; global rebinds at module scope.
Decorators
A decorator is a function taking a function and returning a replacement. @d above def f is exactly f = d(f), evaluated at definition time.
The one obligation is preserving metadata. functools.wraps copies __name__, __doc__, __module__, __qualname__ and __wrapped__ onto the wrapper, without which every decorated function in your codebase reports itself as wrapper in tracebacks and documentation.
Decorators earn their place in ML code for cross-cutting concerns: @lru_cache for expensive pure lookups, @torch.no_grad() for inference, @functools.singledispatch for type-based dispatch, @dataclass for config objects, and timing or retry wrappers around data loading.
Purity, and why it matters here
A pure function depends only on its arguments and mutates nothing outside itself. Purity is not an aesthetic preference in machine-learning code; it is what makes four practical things possible.
- Reproducibility. A preprocessing step that reads a global or mutates its input cannot be rerun to the same result.
- Caching. Memoisation is sound only for pure functions.
- Parallelism.
multiprocessingandjoblibserialise arguments across process boundaries, so hidden shared state either breaks or silently diverges. - Testing. A pure function needs no fixtures.
The two impure things that matter most in practice are global random state and in-place mutation of arrays. Pass an explicit numpy.random.Generator rather than calling np.random.seed, and prefer returning a new array over writing through a view whose owner you cannot see.
Typing
PEP 484 annotations on functions are the highest-value place to use them, because the signature is the contract. PEP 604 allows int | None instead of Optional[int]. typing.Callable, ParamSpec (PEP 612) and TypeVar let a decorator preserve its wrapped signature for a checker.
Annotations are not enforced at run time. pydantic does enforce them, at a real cost, which is why it appears at the edges of a system — parsing configs and API payloads — rather than in inner loops.
References
- PEP 3102 — Keyword-Only Arguments, Talin, 2006.
- PEP 3104 — Access to Names in Outer Scopes, Yee, 2006.
- PEP 318 — Decorators for Functions and Methods, Smith, Coghlan et al., 2003.
- PEP 570 — Python Positional-Only Parameters, Hastings, Galindo et al., 2018.
- PEP 612 — Parameter Specification Variables, Lee, 2019.
- PEP 659 — Specializing Adaptive Interpreter, Shannon, 2021.
- CPython source,
Objects/funcobject.candPython/ceval.c.
What to learn next
- NumPy — ufuncs, and functions that dispatch to compiled kernels.
- Python for machine learning — composing pure stages into a pipeline.
- Loops and conditions — the call-overhead boundary in more detail.