Function calling and tools
Function calling lets a model ask your program to run a specific function with specific arguments, so it can use a calculator, a database or an API instead of guessing.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Function calling is the model writing a request for your code to run a function, and then using the result.
The analogy you have already lived
A doctor examines you and writes a prescription. The doctor does not open a cupboard and hand you tablets. They write a slip: this medicine, this dose, twice a day.
You take the slip to the pharmacist. The pharmacist reads it and gives you the medicine. Then you go back and the doctor uses that in your treatment.
A model does exactly this. It cannot run anything. It writes a slip saying which function to call and with what values. Your program is the pharmacist.
The one thing everyone gets wrong
Read this twice, because almost every beginner gets it backwards.
The model never runs your code. It has no hands. It produces text, and that text happens to be a small structured request.
Your program reads that request, decides whether to honour it, runs the function, and hands the result back. Every decision to actually do something is yours.
That is not a limitation. It is the safety property that makes the whole idea usable.
Why it exists
A language model is a text predictor, and there are things text prediction is bad at.
Arithmetic. Ask for 4,738 times 921 and it will produce a confident, wrong number. It is pattern-matching on what such answers look like.
Today. It has no clock. It does not know the date, the weather, or whether your parcel shipped.
Your data. Your order table, your college's timetable and your bank balance were never in its training.
Doing things. Sending an email, booking a slot, updating a record — a text predictor cannot act on the world.
Function calling connects the model to real code for all four. The model decides what is needed. Ordinary programs do it.
How it works
you: "How much is a dosa, and what's 47 times 89?"
|
v
[ model ] does not calculate, does not look anything up
|
v
writes two slips:
{ name: "get_price", args: { item: "dosa" } }
{ name: "multiply", args: { a: 47, b: 89 } }
|
v
YOUR PROGRAM checks the slips and runs the functions
|
v
results go back: 70 and 4183
|
v
[ model ] writes the final sentence using those numbers
|
v
"A dosa is 70 rupees, and 47 times 89 is 4183."Two trips to the model, with your code doing real work in between. That round trip is the whole technique.
How the model knows what is available
You tell it. Along with the question, you send a short description of each function. That description covers its name, what it does, and what values it needs.
The description is the entire contract. If you write "gets the price" the model will guess when to use it. If you write "Price of one menu item in rupees. Use for food items only." it will guess much better.
Writing those descriptions well is a real skill. It is closer to briefing a new intern than to programming.
Where you have already met it
- An assistant that actually sets your alarm, instead of describing how to.
- A chatbot that quotes your real order status.
- An AI that draws a chart from a spreadsheet you uploaded.
- Anything that says "searching the web..." before answering.
The honest part
The model gets it wrong sometimes. It will ask for a function that does not exist. It will pass the city name where a date was expected. It will call the same function four times in a row.
Small models are noticeably worse at this. The Developer tab shows a real run on a 1-billion-parameter model. It returned the arguments in the wrong shape entirely — a copy of the description instead of the values.
This is not a reason to avoid the technique. It is a reason your code must check every request before running it. Never let a name that came from a model choose which function to run. Look it up in a list you wrote yourself.
Actions that cannot be undone need a human. Reading data is fine to automate. Sending money, deleting records and emailing customers should stop and ask a person first. Build that in from the start, not after the incident.
Remember this
- The model writes a request; your program decides whether to run it.
- You describe the available functions, and the description is the contract.
- Always validate before running, and always ask a human before anything irreversible.
What to learn next
- Structured output — making the model's JSON valid every time.
- AI agents — looping tool calls until a task is done.
- Build an AI agent — put the dispatcher to work on something real.
Developer — Code and libraries.
Setup
Nothing to install for the first example. It is standard library only, and it is the part you actually have to write.
python3 --versionFor the live example you need Ollama running, with a model whose ollama show output lists tools under Capabilities.
The dispatcher — this is the code you own
Every provider's API differs at the edges. The middle is always the same: match a name to a function, check the arguments, run it, wrap the result. Get this right and the provider becomes a detail.
import json
# --- 1. The real Python functions. Ordinary code, nothing AI about them. -------
PRICES = {"idli": 40, "dosa": 70, "filter coffee": 25}
def get_price(item: str) -> dict:
return {"item": item, "rupees": PRICES.get(item.lower(), 0)}
def add(a: float, b: float) -> dict:
return {"sum": a + b}
TOOLS = {"get_price": get_price, "add": add}
# --- 2. Pretend the model replied with this. A real reply looks identical. ----
model_reply = {
"role": "assistant",
"content": "",
"tool_calls": [
{"id": "call_1", "function": {"name": "get_price", "arguments": {"item": "dosa"}}},
{"id": "call_2", "function": {"name": "get_price", "arguments": {"item": "filter coffee"}}},
],
}
# --- 3. The dispatcher. This is the part you actually write. ------------------
def run_tool_calls(calls):
results = []
for call in calls:
name = call["function"]["name"]
args = call["function"]["arguments"]
if name not in TOOLS: # never getattr() a model-supplied name
output = {"error": f"unknown tool {name!r}"}
else:
try:
output = TOOLS[name](**args)
except TypeError as exc: # wrong or missing arguments
output = {"error": str(exc)}
results.append({"role": "tool", "tool_call_id": call["id"],
"name": name, "content": json.dumps(output)})
return results
for message in run_tool_calls(model_reply["tool_calls"]):
print(message)
print()
print("what a bad call produces:")
bad = [{"id": "call_9", "function": {"name": "os.system", "arguments": {"cmd": "rm -rf /"}}},
{"id": "call_10", "function": {"name": "add", "arguments": {"a": 1}}}]
for message in run_tool_calls(bad):
print(message){'role': 'tool', 'tool_call_id': 'call_1', 'name': 'get_price', 'content': '{"item": "dosa", "rupees": 70}'}
{'role': 'tool', 'tool_call_id': 'call_2', 'name': 'get_price', 'content': '{"item": "filter coffee", "rupees": 25}'}
what a bad call produces:
{'role': 'tool', 'tool_call_id': 'call_9', 'name': 'os.system', 'content': '{"error": "unknown tool \'os.system\'"}'}
{'role': 'tool', 'tool_call_id': 'call_10', 'name': 'add', 'content': '{"error": "add() missing 1 required positional argument: \'b\'"}'}The three lines that matter
if name not in TOOLS. A dictionary you wrote is an allowlist. The alternative — getattr(module, name) or eval(name) — hands control of your process to text a model produced. The os.system line in the test is not paranoia; it is the exact shape of the attack.
except TypeError. Models omit required arguments regularly. Catching it turns a crash into a message the model can read and correct on the next turn. Self-repair only works if the error survives back to the model.
Returning an error as a normal result. Notice both failures produced a role: tool message, not an exception. The conversation continues. The model sees "unknown tool" and tries something else. Raising instead would end the request.
The tool description you send
The model needs to know what exists. The format below is JSON Schema, and OpenAI, Anthropic, Google and Ollama all accept a close variant of it. It is a plain data structure, so running the file prints nothing.
TOOL_SCHEMAS = [
{
"type": "function",
"function": {
"name": "get_price",
"description": "Price of one menu item in rupees. Food and drink only.",
"parameters": {
"type": "object",
"properties": {
"item": {"type": "string", "description": "Menu item name, lowercase"},
},
"required": ["item"],
},
},
},
]Every word in description is prompt text. It is read by the model on every request and it drives the decision. Vague descriptions produce vague tool selection, and no amount of code fixes that.
Against a real model
import json, urllib.request
def chat(messages, tools):
body = {"model": "llama3.2:1b", "messages": messages, "tools": tools,
"stream": False, "options": {"temperature": 0}}
request = urllib.request.Request(
"http://localhost:11434/api/chat",
data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(request) as response:
return json.loads(response.read())["message"]
messages = [{"role": "user", "content": "What is the weather in Nagpur right now?"}]
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city.",
"parameters": {"type": "object",
"properties": {"city": {"type": "string", "description": "City name"}},
"required": ["city"]},
},
}]
print(json.dumps(chat(messages, tools), indent=2)){
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "call_lb32pipb",
"function": {
"index": 0,
"name": "get_weather",
"arguments": {
"city": "Nagpur"
}
}
}
]
}The id is generated per call, so yours will differ. content is empty because the model chose to call a tool instead of writing prose — always check tool_calls before checking content.
What failure actually looks like
On a different question, the same 1B model returned this:
"arguments": {
"type": "object",
"required": ["item"],
"properties": {"item": "dosa"}
}It echoed the schema back instead of filling it in. The value dosa is in there, buried one level too deep. Calling get_price(**arguments) raises:
TypeError: get_price() got an unexpected keyword argument 'type'
This is a real observation from a real run, not a hypothetical. Small models do this often enough that you must plan for it. Two defences:
- Validate arguments against your schema before dispatch, with
jsonschemaor a Pydantic model. - Use a model actually trained for tools.
ollama show <model>liststoolsunder Capabilities when it was. Even then, a 1B model is meaningfully less reliable than a 7B one.
Completing the loop
One tool call is half the job. The model needs the result back.
messages.append(assistant_message) # the model's request, unchanged
for result in run_tool_calls(assistant_message["tool_calls"]):
messages.append(result) # one message per call
final = chat(messages, tools) # second round trip
print(final["content"])Three rules that are easy to get wrong.
Append the assistant message exactly as received. Editing it breaks the link between request and result on some providers.
Return one result per call, matched by tool_call_id. Two calls, two results.
Loop with a hard limit. for _ in range(5) around the request-and-dispatch cycle. Models get stuck calling the same tool repeatedly, and an unbounded loop with a paid API is an expensive night.
Common mistakes
Executing before validating. Covered above and worth repeating: allowlist the name, check the argument types, then run.
Automating irreversible actions. Reading is safe to automate. Sending, deleting, paying and posting need a human confirmation step. Split your tools into read and write, and gate the write ones.
Registering twenty tools. Selection accuracy falls as the list grows, and every description is spent context on every request. Keep it under about ten, or route to a smaller subset first.
Returning huge results. A tool that returns 50,000 characters of JSON pushes everything else out of the context window. Summarise or truncate inside the tool.
Assuming JSON arguments are typed. Models return "5" where you expected 5, and "true" where you expected True. Coerce explicitly.
Trusting tool output blindly. If a tool fetches a web page, that page's text now enters your prompt. Text from the internet can contain instructions aimed at your model. Treat every tool result as untrusted input, not as commands.
Try it yourself
Add a third tool to dispatch.py called place_order, which pretends to charge money. Make the dispatcher refuse to run any tool whose name is in a NEEDS_APPROVAL set, returning {"error": "awaiting human approval"} instead.
Then add validation: before calling, check that the arguments dictionary has exactly the keys your schema declares. Feed it the malformed arguments from the failure above and confirm your check catches it.
What to learn next
- Structured output — making the model's JSON valid every time.
- AI agents — looping tool calls until a task is done.
- Build an AI agent — put the dispatcher to work on something real.
Researcher — Mathematics and papers.
What is actually happening at the model level
Function calling is constrained generation plus a fine-tuned convention. Nothing about the transformer changes.
Instruction-tuned models are trained on conversations where the assistant turn emits a special token sequence followed by a structured payload naming a function and its arguments. Llama 3.1 uses <|python_tag|> and reserved role tokens; other families use different delimiters. The serving layer parses that payload and surfaces it as a tool_calls field.
Two consequences follow, and both are routinely misunderstood.
Tool schemas are injected into the prompt. They are rendered into the system turn by the chat template, and they consume context on every request. Twenty tools with long descriptions can cost more than a thousand tokens per call, on every call, whether or not any tool is used.
Reliability is a property of the fine-tune, not the architecture. A model not trained on this format will produce something that looks approximately right and parses badly — which is precisely the failure documented in the Developer tab.
Grammar-constrained decoding
Prompting alone gives no syntactic guarantee. Constrained decoding does, by masking logits at each step to only those tokens that can continue a valid string under a grammar compiled from the schema.
The technique appears as GBNF in llama.cpp, as Outlines (Willard and Louf, 2023), and as XGrammar and LM Format Enforcer. Outlines compiles a regular expression or JSON Schema to a finite-state machine and precomputes, per FSM state, the set of allowable next tokens — reducing per-step overhead to a mask lookup.
Two limits are worth knowing:
- Validity is not correctness. A grammar guarantees a parseable object with the right field types. It cannot make the values true. See structured output.
- Constraint can degrade content quality. Forcing a token the model assigned low probability shifts the conditional distribution for everything after it. Reported effects range from negligible to material depending on how far the constraint pulls from the unconstrained mode.
Parallel and sequential calls
Modern APIs return an array of tool calls per turn. Two distinct execution semantics follow:
- Parallel — calls in one array are independent. Run them concurrently; latency is the max, not the sum.
- Sequential — the model emits one call, sees its result, and only then decides the next. Required when argument values depend on prior results.
This is the ReAct loop (Yao et al., 2022): interleaved reasoning traces and actions, with observations fed back. The formal object is a policy over a state that grows by one observation per step, and the practical constraint is that the state is the context window. Long tool-use trajectories exhaust it, which is why summarisation of older turns becomes necessary.
The security model
Treat every tool result as untrusted input. This is not a general caution; it is a specific, demonstrated attack class.
Indirect prompt injection (Greshake et al., 2023): a document, web page or email that a tool retrieves contains text addressed to the model. Because retrieved content and user instructions occupy the same context with no privilege boundary, the model may follow it. Combined with a tool that can exfiltrate — send an email, make a request to an attacker's URL — this is a complete data-theft chain.
There is no prompt-level fix. Instruction-versus-data separation is not enforceable inside a single token stream. The mitigations are architectural:
- Capability restriction. Tools receive the narrowest scope that works. A read-only database role, not an admin one.
- Human confirmation for irreversible actions. The dual-LLM and action-approval patterns formalise this.
- Provenance tracking. Mark which context regions came from untrusted sources and forbid privileged tool calls on turns that consumed them.
- Egress control. Constrain outbound network destinations so exfiltration has nowhere to go.
Willison's framing is worth internalising: any system combining private data access, untrusted content exposure, and external communication has an open exfiltration path unless one of the three is removed.
Evaluation
Tool selection is a classification problem with an abstention class, and it should be measured as one.
- Selection accuracy — correct tool chosen, given a ground-truth label.
- Argument accuracy — exact match on values, and schema-validity rate separately. These fail differently.
- Abstention — when no tool applies, does the model answer directly instead of forcing a call? Precision on this is commonly the weakest number in a real system.
- Trajectory success — for multi-step tasks, did the sequence reach the goal within a step budget?
Benchmarks: the Berkeley Function Calling Leaderboard (Patil et al.) covers simple, parallel, multiple and irrelevance-detection categories; τ-bench (Yao et al., 2024) evaluates multi-turn tool use against a rule-following user simulator. Both suffer the standard contamination and prompt-sensitivity caveats, so treat them as a shortlist filter and evaluate on your own tool set.
Model Context Protocol
MCP (Anthropic, 2024) standardises the transport and discovery layer rather than the calling convention. A server exposes tools, resources and prompts over JSON-RPC; a client discovers and invokes them. It turns an N-clients × M-tools integration problem into N + M.
The security analysis above applies unchanged, and gains a supply-chain dimension: a third-party MCP server's tool descriptions are prompt text your model reads. Audit them as you would audit a dependency.
References
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, 2022 — arxiv.org/abs/2210.03629
- Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools, 2023 — arxiv.org/abs/2302.04761
- Patil et al., Gorilla: Large Language Model Connected with Massive APIs, 2023 — arxiv.org/abs/2305.15334
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, 2023 — arxiv.org/abs/2302.12173
- Willard and Louf, Efficient Guided Generation for Large Language Models (Outlines), 2023 — arxiv.org/abs/2307.09702
- Yao et al., τ-bench: A Benchmark for Tool-Agent-User Interaction, 2024 — arxiv.org/abs/2406.12045
What to learn next
- Structured output — making the model's JSON valid every time.
- AI agents — looping tool calls until a task is done.
- Build an AI agent — put the dispatcher to work on something real.