smolagents — code-acting agents
smolagents is a small Hugging Face library where the agent acts by writing short Python programs instead of filling in one tool call at a time.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
smolagents is a small Python library where the agent acts by writing and running little pieces of code.
Picture two students solving a long maths problem. The first raises a hand after every single step and waits for the teacher to approve it. The second writes the full working on paper, checks the result, and hands in the answer.
Most agent frameworks work like the first student. smolagents bets on the second one.
Why it exists
An agent is a program that lets a model decide which actions to take — we cover the idea itself in AI agents. The usual way to act is a tool call: the model fills in a small form naming one function and its inputs. That works, and function calling explains it.
But many real tasks need loops, comparisons, and combining results. Filling one form per step makes that slow and clumsy. Ten small steps means ten round trips to the model.
Writing code handles all of that in one go. Code can loop, branch, do arithmetic, and pass one tool's output into another. So the Hugging Face team built an agent that acts by writing Python. They also kept the library small on purpose — the core is short enough to read in an afternoon.
How it works
your task
|
v
[ model writes a short Python snippet ]
|
v
[ smolagents runs it in a fenced-off sandbox ]
|
v
result or error ──→ shown back to the model ──→ it tries again
|
v
the model calls final_answer(...) and the run endsThe loop repeats until the model calls a special function named final_answer, or runs out of allowed steps.
The word sandbox means a fenced play area. The generated code runs inside a restricted mini-interpreter that blocks most imports and dangerous operations. The model cannot delete your files because the sandbox refuses to run that code.
Where you have already seen this
When a chat assistant analyses a spreadsheet you uploaded, it works this way. It writes a small piece of code, runs it, reads the numbers that came back, and writes the next piece. You see a polished answer; underneath was a code-run-look loop.
The honest part
Letting a model run code on your machine is a real risk, and the smolagents authors say so themselves. The built-in sandbox blocks the obvious attacks, but for anything serious the code should run inside a separate, throwaway environment. Treat this the way you treat a pressure cooker: very useful, and worth respecting.
Remember this
- smolagents agents act by writing Python, not by filling one tool-call form per step.
- Code can loop and combine tools, so tasks finish in fewer round trips.
- The generated code runs in a sandbox, because model-written code is untrusted code.
What to learn next
- Designing good tools for agents — the craft that outlives every framework.
- Orchestrating multiple agents — where
managed_agentsfits. - Ollama — the local model server used in this lesson.
Developer — Code and libraries.
This lesson is written against smolagents 1.26 (Python 3.10 or newer). The library moves fast, so pin the version.
Setup
pip install "smolagents==1.26.0"The base install pulls a handful of small dependencies and no model weights. The download is a few tens of megabytes.
A tool the agent can use — runnable with no model at all
Everything below runs offline and prints the same output on every machine.
from smolagents import tool
# Fixed demo rates: rupees per one unit of each currency.
RATES = {"USD": 83.2, "EUR": 90.1, "GBP": 105.4}
@tool
def to_inr(amount: float, currency: str) -> float:
"""
Converts an amount of foreign money into Indian rupees.
Args:
amount: How much money to convert.
currency: A three-letter code. One of USD, EUR or GBP.
"""
if currency not in RATES:
raise ValueError(f"Unknown currency {currency}. Use one of: {sorted(RATES)}")
return round(amount * RATES[currency], 2)
print(to_inr(amount=100, currency="USD"))
print(to_inr.name)
print(sorted(to_inr.inputs))8320.0 to_inr ['amount', 'currency']
What @tool actually did
The decorator turned an ordinary function into a Tool object. It read three things and refuses to work without them:
- the type hints (
amount: float,currency: str), which become the input types the model sees, - the first line of the docstring, which becomes the tool description,
- the
Args:section, which describes each input.
All of that gets baked into the agent's system prompt, exactly as described in function calling. The docstring is not decoration here. It is the tool's user manual, and the model is the user.
The ValueError branch matters more than it looks. When the agent's code hits it, the message is shown back to the model, which can read it and retry with a valid code. Error messages are part of your interface — designing good tools for agents is a whole lesson on this.
Running the agent on a local model
The agent itself needs an LLM. The cheapest honest option is a local one through Ollama:
pip install "smolagents[litellm]==1.26.0"
ollama pull llama3.2 # about a 2 GB download, one timefrom smolagents import CodeAgent, LiteLLMModel
from currency_tool import to_inr
model = LiteLLMModel(
model_id="ollama_chat/llama3.2",
api_base="http://localhost:11434",
num_ctx=8192, # Ollama's default of 2048 truncates agent prompts
)
agent = CodeAgent(tools=[to_inr], model=model, max_steps=4)
print(agent.run("Use the to_inr tool to convert 250 EUR to rupees."))There is no output block here, deliberately. The trace depends on the model, its version, and the run. What you should see: a step where the model writes something like result = to_inr(amount=250, currency="EUR"), the sandbox printing 22525.0, and a final step calling final_answer. A small model will sometimes take a wrong turn first — that is normal, and watching it recover is the best free lesson in agent behaviour.
Common mistakes
Skipping the Args: descriptions. The decorator rejects the function immediately:
smolagents._function_type_hints_utils.DocstringParsingException: Cannot generate JSON schema for to_inr because the docstring has no description for the argument 'amount'
The fix is to describe every argument. This strictness is a feature — a tool the model cannot understand is worse than no tool.
Skipping type hints. Same story, different error:
smolagents._function_type_hints_utils.TypeHintParsingException: Argument amount is missing a type hint in function to_inr
Leaving Ollama's context at its default. With num_ctx at 2048, the agent's system prompt gets cut off mid-sentence and the model behaves bizarrely. Set 8192 or more, as above.
Authorising imports casually. CodeAgent(..., additional_authorized_imports=["requests"]) lets model-written code make network calls. Add imports one at a time, and for anything untrusted run the code remotely with executor_type="docker" or executor_type="e2b" instead of locally.
Try it yourself
Add a second tool from_inr(amount, currency) that converts the other way. Then ask the agent: "Convert 5000 rupees to USD, then convert the result back, and report any difference." Watch how one code snippet chains both tools — the thing a form-filling agent would need four round trips for.
What to learn next
- Designing good tools for agents — the craft that outlives every framework.
- Orchestrating multiple agents — where
managed_agentsfits. - Ollama — the local model server used in this lesson.
Researcher — Mathematics and papers.
The paradigm: code as the action space
smolagents implements the CodeAct idea. Wang et al. (2024), Executable Code Actions Elicit Better LLM Agents (arxiv.org/abs/2402.01030), consolidate the agent's action space into executable Python rather than per-step JSON. Against JSON and text-action baselines, they report up to 20% higher success on API-Bank and their curated M³ToolEval, with fewer actions per solved task.
The mechanism behind the gain is expressiveness. A JSON tool call is one edge in the interaction graph; a code block is a small program over the tool set — composition, control flow, and intermediate state come for free. The lineage runs back to Yao et al. (2022), ReAct (arxiv.org/abs/2210.03629), which interleaved reasoning traces with atomic actions.
Cost structure
Each agent step is a full-context model call: the system prompt, tool schemas, and the entire step history are re-sent every time. For a run of $n$ steps with mean step-transcript length $c$ tokens, input-token cost grows as $\sum_{i=1}^{n} i \cdot c = O(n^2 c)$ — quadratic in steps. Symbols: $n$ is the number of loop iterations, $c$ the tokens appended per iteration.
This is the quantitative argument for code actions: if writing code cuts a 10-step trajectory to 3 steps, it cuts far more than 70% of input tokens. It is also the argument for prompt caching where the provider supports it.
Execution security
The local executor is a restricted interpreter over the Python AST, not a bare exec. Defaults as of 1.26: an import allowlist, submodule access denied unless explicitly authorised (numpy does not grant numpy.random), a cap on iterations, and no attribute access into dunder machinery. Remote executors (Docker, E2B, Blaxel) move execution out of process entirely.
The threat model worth taking seriously is not the model "going rogue" but prompt injection: a tool that fetches web content can return text that instructs the model, and the model may then write code serving the attacker. Sandboxing bounds the damage; it does not prevent the manipulation. Treat any agent with both untrusted input and consequential tools as a security boundary problem.
Multi-agent composition and benchmarks
An agent given name and description attributes can be passed as managed_agents to a manager agent — the same schema treatment as a tool, one level up. The Hugging Face team used this structure for their GAIA submission. Mialon et al. (2023), GAIA (arxiv.org/abs/2311.12983), remains the reference benchmark for tool-using generalist assistants: humans score 92%, while GPT-4 with plugins scored 15% at publication — a gap that made agent scaffolding, not raw model quality, the differentiator.
State of the ecosystem
As of mid-2026, smolagents sits at 1.26 and the API surface still shifts between minor versions — InferenceClientModel, LiteLLMModel, TransformersModel, and the OpenAI/Azure/Bedrock model classes are current names. Pin exact versions in anything you ship, and re-read the release notes before upgrading. The library's differentiators against LangGraph, CrewAI and AutoGen are the code-action default and the deliberately small core, not feature count.
What to learn next
- Designing good tools for agents — the craft that outlives every framework.
- Orchestrating multiple agents — where
managed_agentsfits. - Ollama — the local model server used in this lesson.