AutoGen — conversational multi-agent systems
AutoGen builds AI systems as a group chat — agents talk to each other in turns until a stop condition ends the conversation.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
AutoGen is a framework where AI agents work by talking to each other in a group chat, and the conversation itself is the program.
Think of a classroom group discussion. Students sit in a circle and speak in turns — one proposes an idea, the next criticises it, the next improves it. The teacher does not script what anyone says. The teacher only decides who speaks next and when the discussion is over.
That is AutoGen's whole design. You create the participants, put them in a chat, and set the stopping rule. The solution emerges from the exchange of messages.
Why it exists
Microsoft Research noticed something in 2023: two models talking often solve what one model cannot. A writer improves when a critic answers back. A coder improves when a tester reports what broke.
But wiring up "model A's output becomes model B's input, then back again, until done" by hand is messy plumbing. AutoGen turned the chat itself into the building block. Instead of drawing flowcharts, you configure a conversation.
How it works
Three ideas carry everything.
- An agent — a participant with a name, instructions, and a model behind it.
- A team — the group chat, plus the rule for whose turn it is. The standard one goes around the circle in order.
- A termination condition — the explicit stopping rule: after so many messages, or when someone says a magic word like "APPROVED".
task: "propose a lesson idea"
│
▼
┌──────────────────────────────────┐
│ group chat │
│ │
│ writer ──► editor ──► writer │ turns go around
│ ▲ │ │ the circle...
│ └─────────────────────┘ │
└──────────────────────────────────┘
│
stop rule fires (e.g. 5 messages)
▼
final transcriptThe answer is the transcript. You read the last message, or the whole exchange.
Where you have already seen this shape
- A WhatsApp group deciding a wedding menu: proposal, objections, revision, agreement.
- A code review: author posts, reviewer comments, author fixes, reviewer approves.
- Bargaining at a market stall: two parties, alternating turns, a stopping point.
The honest part, part one
Conversations without a firm stopping rule ramble — models thank each other, repeat themselves, and burn money. In AutoGen the stopping rule is not a detail. It is half the engineering.
The honest part, part two
AutoGen's maker has moved on. In late 2025 Microsoft announced the Microsoft Agent Framework as its successor, and AutoGen is now in maintenance mode — kept working, community-managed, but not growing. It is still worth understanding: real codebases use it, interviews ask about it, and its core idea — conversation as the program — shaped every framework in this section. Learn the idea; hold the specific library loosely.
Remember this
- AutoGen agents solve tasks by talking in a group chat, taking turns.
- The termination condition is an explicit, first-class part of the design.
- AutoGen is in maintenance mode with a Microsoft-built successor — learn its concepts, and check the project's status before starting anything new on it.
What to learn next
- MCP — the Model Context Protocol — the tool-connection standard every framework here now speaks.
- CrewAI — role-based agent teams — task pipelines instead of open conversation.
- LangGraph — agents as graphs — explicit control flow, the opposite design bet.
Developer — Code and libraries.
Setup
pip install autogen-agentchat autogen-extWritten against autogen-agentchat 0.7. Careful with names: the modern packages are autogen-agentchat (the high-level chat API), autogen-core (the runtime) and autogen-ext (model clients and tools). The old package pyautogen and tutorials built on ConversableAgent belong to the 0.2 generation — a different API that will not match anything below. The install is modest by agent-framework standards, in the tens of megabytes.
AutoGen is async-first: everything runs inside asyncio. And like the other lessons in this section, the example uses a scripted model client — ReplayChatCompletionClient, which AutoGen itself provides for testing — so it runs offline, free, and identically on your machine.
A writer-editor team you can run
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.conditions import MaxMessageTermination
from autogen_ext.models.replay import ReplayChatCompletionClient
# a scripted model client: it replays these answers in order,
# so the example runs offline and costs nothing
model = ReplayChatCompletionClient([
"Idea: a lesson on how UPI fraud detection works.",
"Feedback: good idea, but narrow it to one real fraud pattern.",
"Revised: a lesson on how banks catch SIM-swap fraud in seconds.",
"Feedback: approved.",
])
writer = AssistantAgent("writer", model_client=model,
system_message="You suggest lesson ideas.")
editor = AssistantAgent("editor", model_client=model,
system_message="You give blunt feedback on ideas.")
# agents speak in a fixed circle until the stop condition fires
team = RoundRobinGroupChat([writer, editor],
termination_condition=MaxMessageTermination(5))
async def main():
result = await team.run(task="Propose one lesson for an AI learning site.")
for msg in result.messages:
print(f"{msg.source}: {msg.to_text()}")
asyncio.run(main())user: Propose one lesson for an AI learning site. writer: Idea: a lesson on how UPI fraud detection works. editor: Feedback: good idea, but narrow it to one real fraud pattern. writer: Revised: a lesson on how banks catch SIM-swap fraud in seconds. editor: Feedback: approved.
For a real model, swap the client — for example OpenAIChatCompletionClient(model="gpt-4.1") from autogen_ext.models.openai, installed with pip install "autogen-ext[openai]" and an API key. Agents and team stay untouched.
Line by line, the parts that are not obvious
RoundRobinGroupChat([writer, editor], ...) — the list order is the speaking order. The task message counts as the first message, from user; then writer, editor, writer, editor. With the limit at 5, the run stops exactly where you see it stop.
MaxMessageTermination(5) — termination conditions are objects, and they compose. MaxMessageTermination(20) | TextMentionTermination("APPROVED") stops on whichever fires first. That one line — a hard cap OR a success word — is the standard production pattern.
await team.run(task=...) — one call runs the whole conversation to termination and returns a result whose .messages is the transcript. There is also run_stream() for printing messages as they happen, which you will want for anything interactive.
msg.to_text() — messages are typed objects, not strings; .source says who spoke, .to_text() gives the readable content.
Why the transcript matters: the editor's "narrow it" message became part of the writer's context on its next turn — that is the entire mechanism. No routing code exists. Whatever is in the chat is what the next speaker sees.
Common mistakes
Mixing API generations. Copying a ConversableAgent or initiate_chat snippet from an old blog into a 0.7 install fails on the first import. Before debugging anything, check which generation your tutorial is from: pyautogen imports mean 0.2, autogen_agentchat means the modern line.
No termination condition. Without one, the round-robin keeps circling until something else stops it. With a real model that something is your wallet. Set a cap on every team, even in experiments.
Forgetting the event loop. team.run(...) without await, or await outside an async function, gives you a coroutine object or a SyntaxError respectively. All AutoGen code lives under asyncio.run(...) somewhere.
One venv for every framework. Installing this lesson's frameworks together in one environment produces real dependency clashes — while preparing these lessons, autogen-core pinned protobuf to a 5.x version and CrewAI's install pulled in 6.x, breaking one of them. One virtual environment per framework saves an evening of pip archaeology.
Try it yourself
Add a third agent, fact_checker, to the circle, extend the scripted replies so each agent gets a sensible line, and raise the message cap to 7. Then replace MaxMessageTermination(5) with TextMentionTermination("approved") and work out why the run now ends one message earlier than you expected — the answer teaches you exactly when conditions are checked.
What to learn next
- MCP — the Model Context Protocol — the tool-connection standard every framework here now speaks.
- CrewAI — role-based agent teams — task pipelines instead of open conversation.
- LangGraph — agents as graphs — explicit control flow, the opposite design bet.
Researcher — Mathematics and papers.
The paper
AutoGen is one of the few frameworks in this section with a proper citation: Wu et al., 2023, AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation (arxiv.org/abs/2308.08155). Its framing, conversation programming, treats the message sequence as the control flow: agents are "conversable" units with a unified send/receive interface, and computation advances by message exchange rather than function composition. The paper's applications — maths solving, code generation with execution feedback, decision-making — established the writer-critic and coder-executor loops as standard patterns.
The striking design choice against the rest of this section: where LangGraph makes control flow explicit and state central, AutoGen makes both implicit in the transcript. The transcript is the state, and the speaker-selection policy is the control flow.
The 0.4 rewrite: an actor system underneath
The 0.2 → 0.4 rewrite (January 2025) rebuilt AutoGen on an asynchronous, event-driven actor model in autogen-core: agents are actors subscribing to typed message topics, and the runtime routes messages between them — a lineage running back to Hewitt's actor formalism (Hewitt, Bishop and Steiger, 1973) via modern implementations like Orleans. autogen-agentchat reconstructs the friendly group-chat API on top of that substrate. Practical consequences: cross-language agents (Python and .NET share the runtime), distributed deployment, and streaming as the default execution mode.
Speaker selection generalises beyond round-robin: SelectorGroupChat asks a model to choose the next speaker from agent descriptions each turn (one extra model call per turn), and Swarm implements handoff-based transfer where the current speaker names its successor — the pattern popularised by OpenAI's Swarm sketch.
Magentic-One and the orchestrator pattern
Fourney et al., 2024, Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks (arxiv.org/abs/2411.04468) is the flagship system built on this stack: an Orchestrator agent maintains a task ledger (facts, guesses, plan) and a progress ledger, delegating to web-surfer, file-surfer, coder and terminal agents, re-planning when progress stalls. Its evaluation on GAIA, AssistantBench and WebArena made the two-ledger orchestrator a reference architecture — and its error analyses (loops of repeated actions, stale ledger facts) remain the honest catalogue of what still goes wrong.
Maintenance mode and the successor
In October 2025 Microsoft announced the Microsoft Agent Framework, merging AutoGen's research line with Semantic Kernel's enterprise line; the AutoGen repository now states it is in maintenance mode, community-managed, with new feature work happening in the successor. For the reader in 2026 this means: the concepts above are durable, the package is stable but frozen, and greenfield Microsoft-stack work should evaluate the successor first. Verify the current status before committing — this paragraph ages by design.
Where the evidence stands
Multi-agent conversation inherits the mixed evidence covered in the CrewAI lesson: debate helps on some tasks (Du et al., 2023, arxiv.org/abs/2305.14325), much of the gain reduces to extra samples (Li et al., 2024, arxiv.org/abs/2402.05120), and most observed failures are coordination failures, not capability failures (Cemri et al., 2025, arxiv.org/abs/2503.13657). Cost scales with transcript length: a $T$-turn chat whose transcript is re-read each turn costs $O(T^2)$ in tokens — where $T$ is the number of turns — which is the quantitative argument for aggressive termination conditions.
What to learn next
- MCP — the Model Context Protocol — the tool-connection standard every framework here now speaks.
- CrewAI — role-based agent teams — task pipelines instead of open conversation.
- LangGraph — agents as graphs — explicit control flow, the opposite design bet.