Agentic Frameworks

Designing good tools for agents

An agent is only as good as the tools you hand it — this lesson covers naming, schemas, idempotence and error messages that help the model recover.

On this page 7
  1. Why this matters so much
  2. How it works
  3. The four habits of a good tool
  4. Where you have already seen this
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A tool is a button you hand to the model, and the label on that button decides almost everything.

Think about an ATM. Total strangers walk up to it with no training, and they must succeed on the first try. So it has very few buttons, each button has a plain name, and each button does exactly one thing.

And notice what happens when things go wrong. The screen says "insufficient balance" — a message that tells you what to do next. It does not say "error 37".

Why this matters so much

An agent is a model that acts through tools — functions it can call, as covered in AI agents and function calling.

Here is the part beginners miss. The model never sees your code. It sees only the tool's name, its description, and the names of its inputs. That tiny label is the entire interface.

So when an agent picks the wrong tool, or calls the right tool wrongly, the cause is usually not a weak model. The cause is a confusing toolbox. Tool design is where agent quality is actually decided, and it will still matter when today's frameworks are forgotten.

How it works

      what the model sees              what actually runs
 ┌────────────────────────────┐      ┌───────────────────┐
 │ name:  cancel_order        │      │ your Python code, │
 │ does:  cancels one order   │  ──→ │ your database,    │
 │ needs: order_id            │      │ your API calls    │
 └────────────────────────────┘      └───────────────────┘
        the label                        the machinery

The model reads labels, chooses one, and fills in the inputs. Your machinery runs. Whatever text you return becomes the model's next clue.

The four habits of a good tool

  • One job. "Cancel an order" is a tool. "Manage orders" is a maze.
  • A name that says the job. cancel_order, not process_request.
  • Safe to press twice. Checking a balance twice is harmless. Withdrawing twice is a disaster. Build tools so an accidental repeat does no extra damage — this property is called idempotence, and you will meet it again below.
  • Errors that coach. Every failure message should tell the model what to try instead.

Where you have already seen this

The UPI app on your phone is a masterclass. One clear Pay button. A PIN gate before anything irreversible. And when a payment fails, the message names the reason — wrong PIN, limit exceeded, bank server down — so you know your next move.

The honest part

Naming is genuinely hard, and nobody gets it right on the first pass. Real tool design is a loop: watch transcripts of the agent working, find the tool it misused, sharpen the name or description, watch again. Budget for that loop.

Remember this

  • The model sees only the label — name, description, inputs. Perfect that label.
  • One tool, one job, safe to call twice.
  • An error message is advice to the model, not a complaint into the void.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version   # standard library only — nothing to install

Deliberately framework-free. These rules apply identically in smolagents, LangGraph, CrewAI, or a raw API call — that is what makes them worth learning.

The label: a schema

This is what the model receives, in the JSON-schema style every major provider uses (see function calling for how it is sent):

json
{
  "name": "cancel_order",
  "description": "Cancel one unshipped order. Safe to call twice with the same id.",
  "parameters": {
    "type": "object",
    "properties": {
      "order_id": {
        "type": "string",
        "description": "The order id, e.g. 'OD-1002'. Get ids from list_orders."
      }
    },
    "required": ["order_id"]
  }
}

Three deliberate choices are hiding in there. The description says what the tool is for and one safety fact. The parameter description shows an example value. And it names the sibling tool (list_orders) that produces valid inputs — you are teaching the model the workflow, not only the function.

The machinery: every branch coaches the model

tool_contract.py
import json

ORDERS = {
    "OD-1001": {"item": "USB cable", "status": "shipped"},
    "OD-1002": {"item": "keyboard", "status": "processing"},
}

def cancel_order(order_id: str) -> str:
    # Every branch returns advice the model can act on, never a bare failure.
    if order_id not in ORDERS:
        return json.dumps({"error": f"No order '{order_id}'. Ids look like 'OD-1001'. Call list_orders first."})
    order = ORDERS[order_id]
    if order["status"] == "cancelled":
        return json.dumps({"ok": True, "order_id": order_id, "note": "was already cancelled, nothing changed"})
    if order["status"] == "shipped":
        return json.dumps({"error": "Already shipped, cannot cancel. Offer the customer a return instead."})
    order["status"] = "cancelled"
    return json.dumps({"ok": True, "order_id": order_id, "note": "cancelled"})

print(cancel_order("OD-9999"))
print(cancel_order("OD-1001"))
print(cancel_order("OD-1002"))
print(cancel_order("OD-1002"))   # the retry — same call, same calm answer
Output
{"error": "No order 'OD-9999'. Ids look like 'OD-1001'. Call list_orders first."}
{"error": "Already shipped, cannot cancel. Offer the customer a return instead."}
{"ok": true, "order_id": "OD-1002", "note": "cancelled"}
{"ok": true, "order_id": "OD-1002", "note": "was already cancelled, nothing changed"}

Read those four lines — they are four rules

Line 1: an error with a recovery path. Wrong id. The message shows the id format and names the tool that lists valid ids. An agent reading this can fix itself in one step.

Line 2: an error with an alternative. The rule "shipped orders cannot be cancelled" lives in your system, not in the model's head. The message states the rule and hands the model its next move.

Line 3: a success that names its effect. Small, structured, no padding.

Line 4: idempotence, demonstrated. The repeated call changed nothing and said so. Agents retry — after timeouts, after crashes, after losing track mid-conversation. Idempotent means a repeat call leaves the world exactly as the first call left it. For payments-style tools where a genuine second action is sometimes wanted, add an idempotency_key parameter so retries can be told apart from new requests.

Notice also what the returns are not: they are not raw database dumps. A tool that returns 3,000 tokens of nested JSON spends your context window and buries the one field that mattered.

Common mistakes

Vague names. process_data, handle_request, do_task. If a new teammate could not guess the behaviour from the name, the model cannot either.

Overlapping tools. search_orders, find_orders, and query_orders side by side force the model to gamble. Merge them, or state in each description when it wins over its siblings.

Stack traces as error surface. A raw KeyError: 'OD-9999' teaches the model nothing and sometimes derails it. Catch, translate into advice, return.

Free-text where a fixed choice exists. If status must be one of three values, declare an enum in the schema instead of hoping. The same instinct as structured output.

Too many tools. Past a few dozen, selection quality drops and the schemas eat your prompt. Split responsibilities across specialised agents instead — see orchestrating multiple agents.

Try it yourself

Write the sibling list_orders() tool. Give it a one-line description, make it return at most five orders with only order_id, item, and status, and decide what it should say when there are zero orders. Then show both schemas to a friend and ask them to guess each tool's behaviour. The friend test is the cheapest evaluation you will ever run.

What to learn next

Researcher — Mathematics and papers.

Formal framing

A tool is a typed, partial function $t_i : X_i \rightharpoonup Y_i$ with a natural-language annotation $d_i$ (name, description, parameter docs). The agent's policy $\pi$ selects, at each step, an action from $\mathcal{A} = {(t_i, x)} \cup {\text{respond}}$ conditioned on the transcript. Symbols: $X_i$ is the schema-defined input space, $Y_i$ the output space, $\rightharpoonup$ marks partiality (calls can fail), and $d_i$ is the only information about $t_i$ available to $\pi$.

Two consequences follow. First, tool selection is a discrimination task over ${d_i}$: semantically overlapping annotations raise selection entropy and measured error. Second, the error surface is part of $Y_i$ — a return that encodes corrective information changes the posterior over next actions, which is why coached errors outperform opaque ones in multi-step tasks.

Evidence

  • Schick et al. (2023), Toolformer (arxiv.org/abs/2302.04761) — models can teach themselves API use from self-supervised signals; tool docs function as the supervision channel.
  • Patil et al. (2023), Gorilla (arxiv.org/abs/2305.15334) — grounding calls in retrieved, current API documentation measurably reduces hallucinated APIs and wrong-argument calls.
  • The Berkeley Function Calling Leaderboard (gorilla.cs.berkeley.edu/leaderboard.html) scores call generation by AST-matching, and separately scores relevance detection — declining to call when no tool fits. Irrelevance handling remains a common failure and is directly a function of description quality.
  • Yao et al. (2024), τ-bench (arxiv.org/abs/2406.12045) — agents with identical toolsets diverge widely once domain policy rules interact with tools, and reliability under repetition (pass^k) collapses far below single-run success. Tool contracts that encode policy ("cannot cancel shipped orders") into returns are the mitigation this lesson teaches.
  • Wang et al. (2024), What Are Tools Anyway? (arxiv.org/abs/2403.15452) — a survey unifying tool-use terminology and measuring efficiency trade-offs.

Cost accounting

Schemas are paid for on every call: with $k$ tools averaging $s$ tokens of schema, every agent step carries $ks$ tokens of fixed overhead before any content. Symbols: $k$ tool count, $s$ mean schema size. At $k=40, s=150$ that is 6,000 tokens per step — often more than the conversation itself. Mitigations, in rising complexity: trim descriptions, retrieve a per-task tool subset, split tools across specialist agents, or rely on provider-side prompt caching of the static prefix.

Idempotence, precisely

For a tool with side effects on state $\sigma$, idempotence requires $f(f(\sigma)) = f(\sigma)$ where $f$ is the state transition of one call with fixed arguments. Agent loops give at-least-once execution semantics — retries after timeouts, replays after checkpoint restores — so non-idempotent mutations demand deduplication. The standard mechanism is a client-supplied idempotency key, as in payment APIs: the server stores the key with the first result and replays that result for duplicates. Design reads to be pure, mutations to be idempotent, and anything irreversible to require a distinct confirmation step — which is where human-in-the-loop agents enters.

Distribution

Tool definitions are increasingly shipped as servers rather than in-process functions — the Model Context Protocol standardises the schema-plus-transport layer, as of 2026 the de facto interchange format. The design rules here transfer unchanged: MCP moves tools between processes; it does not make a badly named tool good.

What to learn next