Agentic Frameworks

Orchestrating multiple agents

Splitting one big job across several specialist agents — the patterns that work, the message plumbing between them, and why one good agent often still wins.

On this page 7
  1. Why it exists
  2. The common shapes
  3. How it works
  4. Where you have already seen it
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Multi-agent orchestration means splitting one big job across several agents and managing how they talk to each other.

Think of a wedding kitchen feeding five hundred guests. No single cook makes the biryani, the dosas and the dessert. A head cook takes the order, splits it across stations, and checks every plate at the pass before it goes out.

Each station keeps its own counter, its own tools, its own small mess.

Why it exists

You met single agents in AI agents. Why would anyone want several?

A crowded toolbox confuses the model. Hand one agent forty tools and it starts picking wrong — the selection problem from designing good tools. Two agents with ten relevant tools each choose better.

A crowded conversation confuses it more. An agent researching laptops fills its context window with search results. If the same agent must then write a polished reply, the clutter comes along. A separate writer starts with a clean counter and only the useful notes.

Some work can happen side by side. Three stations cook at once; one cook cannot.

The common shapes

  • Manager and workers — one agent splits the job, others execute, the manager assembles. The wedding kitchen.
  • Assembly line — each agent transforms the work and passes it on. Research, then write, then check.
  • Writer and critic — one produces, one reviews, the work bounces between them a fixed number of times.
  • Handoff — an agent realises the job belongs to a colleague and transfers the whole conversation. A support call moved from sales to billing.

How it works

 customer request
        |
        v
   [ manager ] ── "find laptops under 60k" ──→ [ researcher ]
        ▲                                            |
        |                                          notes
        |                                            v
   final reply ←── [ writer ] ←── "draft a reply from these notes"

The manager never cooks. The researcher never talks to the customer. Each agent sees only its own slice — that separation is the entire trick.

Where you have already seen it

A support ticket that moves from the chat bot to the billing team to the delivery partner is a handoff chain run by people. Coding assistants increasingly run a quiet second agent that reviews the first one's code before you see it — a writer-critic pair.

The honest part

More agents is usually worse. Every extra agent adds cost, delay, a new place for information to get lost, and a harder debugging day. Researchers who studied failing multi-agent systems found most broke through bad orchestration — vague roles, garbled handoffs — not weak models. Start with one good agent. Split only when a single context genuinely cannot hold the job.

Remember this

  • Split for focus: fewer tools per agent, cleaner context per agent.
  • The manager's real job is decomposition and assembly, not doing the work.
  • One good agent beats a committee of confused ones — split late, not early.

What to learn next

Developer — Code and libraries.

Setup

bash
python3 --version   # standard library only — nothing to install

The specialists below are scripted stand-ins, so the file runs identically everywhere with no key and no download. In production an LLM call sits behind each function — the wiring around them is what transfers.

A manager and two workers

kitchen_crew.py
# Two specialists. In a real system an LLM sits behind each function —
# the wiring pattern around them is what this file teaches.
def research_agent(subtask):
    shelf = {"laptops under 60000": "Three fit: Aspire (55k), IdeaPad (58k), Vivobook (59k)."}
    return shelf.get(subtask, "nothing found")

def writer_agent(subtask, notes):
    return f"Hi! {notes} The IdeaPad has the best service network in India."

def manager(request, log):
    # Step 1: decompose. A real manager asks its own LLM to do this split.
    research_task = "laptops under 60000"
    log.append(("manager -> researcher", research_task))
    notes = research_agent(research_task)
    log.append(("researcher -> manager", notes))
    # Step 2: hand only the useful notes to the writer, not the whole history.
    log.append(("manager -> writer", f"draft a reply using: {notes}"))
    reply = writer_agent(request, notes)
    log.append(("writer -> manager", reply))
    return reply

log = []
answer = manager("Suggest a laptop under 60000 rupees", log)
for sender, message in log:
    print(f"{sender:<24}| {message}")
print("---")
print("final:", answer)
Output
manager -> researcher   | laptops under 60000
researcher -> manager   | Three fit: Aspire (55k), IdeaPad (58k), Vivobook (59k).
manager -> writer       | draft a reply using: Three fit: Aspire (55k), IdeaPad (58k), Vivobook (59k).
writer -> manager       | Hi! Three fit: Aspire (55k), IdeaPad (58k), Vivobook (59k). The IdeaPad has the best service network in India.
---
final: Hi! Three fit: Aspire (55k), IdeaPad (58k), Vivobook (59k). The IdeaPad has the best service network in India.

What the wiring teaches

The log is the product. Multi-agent systems are debugged by reading their conversations, the way you debug a team by sitting in its meetings. Build the message log on day one; every framework's tracing UI is a prettier version of this list.

The writer never saw the search. It received distilled notes, not the researcher's whole transcript. That context isolation is the reason to have separate agents at all — pass everything to everyone and you have one bloated agent with extra latency.

The manager owns the joins. Decomposing the request and assembling the result are the manager's whole job, and both are failure points. When you swap the stubs for models, most of your prompt effort goes into the manager's two seams.

Workers are described like tools. The manager routes to "researcher" based on a role description, the same way an agent picks a tool from its description — the naming rules of designing good tools apply verbatim to describing agents.

Moving to real frameworks

As of 2026 every major framework ships this pattern under a different name: CrewAI builds role-based crews, AutoGen frames it as multi-agent conversations, LangGraph wires agents as graph nodes, and smolagents passes worker agents as managed_agents. The concepts in this file map one-to-one onto each; check the framework lesson plus its current docs for exact call names, since these APIs still move.

Common mistakes

Reaching for multi-agent first. The most expensive mistake in the field. A single agent with well-designed tools solves most jobs cheaper, faster, and more debuggably. Split when the context is genuinely crowded, not when the architecture diagram looks nicer.

Sharing full history between agents. It defeats context isolation and multiplies token cost. Pass distilled results across the seams.

No stop condition on agent-to-agent chat. Two polite models can thank each other indefinitely. Cap rounds between any pair, always.

Vague worker descriptions. A manager routing on "helps with various tasks" routes randomly. Describe each worker like a tool: what it does, when to pick it, what to send it.

Evaluating agents one by one but never the system. The seams — decomposition and assembly — are exactly what per-agent tests miss. Run end-to-end tasks through the whole crew, per evaluating agents.

Try it yourself

Add a third specialist, price_checker(model_name), returning a fixed current price for each of the three laptops. Route the manager through it after research, and watch the log grow by two lines. Then add a rounds counter to manager and refuse to exceed six messages — you have implemented the stop condition most demos forget.

What to learn next

Researcher — Mathematics and papers.

Framings in the literature

  • Wu et al. (2023), AutoGen (arxiv.org/abs/2308.08155) — orchestration as programmable multi-agent conversation; agents, humans, and tools share one message-passing abstraction.
  • Li et al. (2023), CAMEL (arxiv.org/abs/2303.17760) — role-play between instructor and assistant agents, an early systematic study of two-agent cooperation and its drift modes.
  • Hong et al. (2023), MetaGPT (arxiv.org/abs/2308.00352) — encodes human standard operating procedures into agent workflows; structured intermediate artefacts, not free chat, cross the seams.
  • Qian et al. (2023), ChatDev (arxiv.org/abs/2307.07924) — a virtual software company as a communicating-agent waterfall.

Does adding agents help?

The affirmative evidence: Du et al. (2023) (arxiv.org/abs/2305.14325) show multi-agent debate improving factuality and reasoning; Li et al. (2024), More Agents Is All You Need (arxiv.org/abs/2402.05120), show sampling-and-voting performance scaling with agent count; Wang et al. (2024), Mixture-of-Agents (arxiv.org/abs/2406.04692), layer aggregators over proposers to beat strong single models on preference benchmarks.

The deflationary reading: much of that gain is ensembling — repeated sampling plus aggregation — which a single model with self-consistency voting approximates at similar token cost. The honest comparison is never "one agent vs five agents" but "equal token budget, differently organised", and papers that control for spend find the gap narrows. Role specialisation earns its keep when tasks decompose along genuinely different tool sets or contexts, which is an engineering property of the task, not a universal law.

Cemri et al. (2025), Why Do Multi-Agent LLM Systems Fail? (arxiv.org/abs/2503.13657), annotate 1,600+ traces across seven frameworks into 14 failure modes in three classes — system design, inter-agent misalignment, and task verification — and find failures dominated by orchestration quality rather than model capability. That is the empirical backing for "split late, not early".

Cost and reliability structure

Token cost: with $a$ agents exchanging $t$ messages of mean context $c$, a shared-transcript topology re-reads history at every hop, giving total input cost $O(a \cdot t^2 \cdot c)$ growth in conversation length; hierarchical topologies passing distilled summaries cut the $t^2$ term back toward linear. Symbols: $a$ agents, $t$ messages, $c$ tokens per message context.

Reliability: a pipeline of $h$ sequential handoffs, each succeeding with probability $r_i$, completes with $\prod_{i=1}^{h} r_i$ — the pass^k arithmetic of evaluating agents applied across the seams, and the quantitative case for keeping $h$ small and each seam checked. Error cascades are the multi-agent-specific failure: a wrong intermediate artefact is consumed downstream as ground truth, so validation at seams (schemas on handoffs, checker steps at the pass) buys more reliability per token than a better model in any single seat.

Topologies and coordination

Star (manager-worker), chain (pipeline), fully connected debate, and blackboard (shared state all agents read and write) are the recurring graphs; termination, deadlock, and livelock — two agents deferring to each other forever — must be handled explicitly, usually by round caps and manager arbitration. Interoperability efforts are consolidating the plumbing: the Model Context Protocol standardises agent-to-tool wiring, and as of 2026 agent-to-agent protocols (Google's Agent2Agent among them) are attempting the same for the seams between agents, though none is yet universal.

Open problems

Credit assignment — which agent caused the failure — lacks principled tooling beyond trace reading. Learned orchestration (a model deciding the topology per task) remains research-grade. And system-level evaluation inherits every difficulty of single-agent evaluation, multiplied by the seams.

What to learn next

What to learn next

These follow on from what you just read.

  • ML Interview Preparation

    The ML interview landscape

    An ML interview loop is four or five separate tests, each checking a different skill — here is what each round asks, and how to divide your preparation.

  • ML Interview Preparation

    ML coding interviews

    The ML coding round asks you to build k-means, attention or backprop from raw NumPy — here is each one, runnable, with the traps interviewers watch for.

  • ML Interview Preparation

    ML theory questions, with answers

    Twenty real ML theory interview questions with worked answers at the depth a strong candidate gives — bias-variance, regularisation, metrics, leakage and more.