CrewAI — role-based agent teams
CrewAI organises AI agents like a work team — each agent gets a role and a goal, each task gets an owner, and work flows from one to the next.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
CrewAI is a framework where you build a team of AI agents, give each one a job title, and hand the team a list of tasks.
Think of a wedding caterer's kitchen. One cook rolls rotis, one stirs the dal, one plates the sweets, and the head cook checks each dish before it leaves. Nobody does everything. Each person has a role, and the food moves down the line from station to station.
CrewAI applies that shape to AI agents. Instead of one model doing everything in one giant prompt, you create several agents, each with a role (its job title), a goal (what it must achieve) and a backstory (a short character sketch that shapes how it behaves).
Why it exists
One model asked to "research this topic, write an article, and edit it" tends to do all three halfway. Long instructions blur together, and the model drifts.
The same model asked to be only an editor — with editing as its entire identity — criticises far more sharply. Splitting one muddled job into several focused jobs is an old idea from human workplaces. CrewAI packages that idea: focused agents, explicit hand-offs.
The other reason is readability. "Researcher passes notes to Writer, Writer passes a draft to Editor" is a system your teammate can understand in one sentence.
How it works
You define three things, and the names mean what they say.
- An agent — role, goal, backstory, and which model powers it.
- A task — a description of one piece of work, plus what a good result looks like, assigned to one agent.
- A crew — the agents plus the tasks plus the running order.
Task 1: "find three ideas" Task 2: "pick the best and improve it"
│ │
▼ ▼
┌────────────┐ ideas list ┌────────────┐
│ Researcher │ ─────────────────► │ Editor │ ──► final answer
└────────────┘ (passed along) └────────────┘Each task's output is handed to the next task as context — like the dish moving to the next station.
Where you have already seen this shape
- A newsroom: reporters gather, sub-editors trim, the editor approves.
- A tailoring shop: one person measures and cuts, another stitches, a third does the finishing.
- Any hospital OPD: nurse takes vitals, doctor diagnoses, pharmacist dispenses.
The honest part
Role-playing teams of agents sounds magical, and demos make it look free. It is not free. Every agent turn is at least one full model call, so a five-agent crew can cost five times what a single good prompt costs — and a well-written single prompt beats a badly designed crew more often than people admit. Reach for a crew when the work genuinely has stages that check each other, not because teams sound impressive.
Remember this
- CrewAI = agents (roles) + tasks (work items) + a crew (the running order).
- Each task's output flows into the next task automatically.
- More agents means more model calls. Split roles when the work has real stages, not for show.
What to learn next
- AutoGen — conversational multi-agent systems — agents that talk in a group chat instead of a task line.
- LangGraph — agents as graphs — the explicit-control-flow alternative.
- Prompt engineering — the skill hiding inside every role and backstory field.
Developer — Code and libraries.
Setup
pip install crewaiWritten against crewai 1.15. Be warned about the size: the install pulls in a vector database, Apache Arrow and an ONNX runtime, among others — several hundred megabytes of dependencies. On the first run, CrewAI prints a one-time question about enabling tracing; answering no saves that preference. It also sends anonymous usage telemetry, which you can switch off by setting the environment variable CREWAI_DISABLE_TELEMETRY=true.
CrewAI normally drives a hosted model (via an API key) or a local one through Ollama. For a lesson, both would make the output depend on your machine and your key. Instead, the example below plugs in a scripted model, which CrewAI supports through its BaseLLM class — so it runs offline, free, and prints the same thing for everyone. The crew mechanics are identical either way.
A two-agent crew you can run
from crewai import Agent, Task, Crew, Process, BaseLLM
# a scripted model, so this runs offline and costs nothing.
# CrewAI accepts any subclass of BaseLLM that implements call().
class ScriptedLLM(BaseLLM):
replies: list[str] = []
calls: int = 0
def call(self, messages, tools=None, callbacks=None,
available_functions=None, from_task=None, from_agent=None,
response_model=None):
reply = self.replies[min(self.calls, len(self.replies) - 1)]
self.calls += 1
return reply
llm = ScriptedLLM(model="scripted", replies=[
"Plan: revise probability basics, then attempt ten past-paper questions.",
"Checked: the plan fits in two hours if each question gets ten minutes.",
])
planner = Agent(
role="Study planner",
goal="Turn a vague study wish into a concrete plan",
backstory="You plan revision sessions for busy students.",
llm=llm,
)
checker = Agent(
role="Plan checker",
goal="Verify a study plan fits the available time",
backstory="You are strict about realistic time budgets.",
llm=llm,
)
plan = Task(
description="Make a two-hour study plan for probability.",
expected_output="A short, concrete plan.",
agent=planner,
)
check = Task(
description="Check the plan fits in two hours.",
expected_output="A one-line verdict.",
agent=checker,
)
crew = Crew(agents=[planner, checker], tasks=[plan, check],
process=Process.sequential)
result = crew.kickoff()
for t in result.tasks_output:
print(f"[{t.agent}]")
print(t.raw)
print("final answer:", result.raw)[Study planner] Plan: revise probability basics, then attempt ten past-paper questions. [Plan checker] Checked: the plan fits in two hours if each question gets ten minutes. final answer: Checked: the plan fits in two hours if each question gets ten minutes.
To use a real model, delete the ScriptedLLM class and build the llm differently — for example llm = LLM(model="ollama/llama3.2", base_url="http://localhost:11434") with LLM imported from crewai, or set an API key and name a hosted model. Nothing else in the file changes.
Line by line, the parts that are not obvious
role, goal, backstory — these are not decoration. CrewAI assembles them into the agent's system prompt. Writing them is prompt engineering wearing a nicer shirt, and vague roles produce vague agents.
expected_output — the most underrated field in the framework. It goes into the prompt as the definition of done. "A one-line verdict" gets you one line; leave it woolly and you get essays.
Process.sequential — tasks run in list order, and each task's output is appended to the context of the next. The alternative, Process.hierarchical, adds a manager model that decides who works next and reviews results — more flexible, and at least one extra model call per hand-off.
result.tasks_output — one entry per task, in order, with .raw holding the text. The crew's result.raw is the final task's output.
Why does ScriptedLLM work? CrewAI talks to models through one interface: call(messages, ...) -> str. Anything honouring that contract is a valid model, which is also how you would wire in a company-internal model gateway.
Common mistakes
No model configured. Without an llm= and without an OPENAI_API_KEY in the environment, the crew fails at kickoff with an authentication error from the default OpenAI client. The fix is explicit: pass llm= to every agent, or set the key.
Treating expected_output as optional paperwork. It is required, and it steers the model harder than the description does. Write it like an acceptance test: "a bulleted list of exactly three ideas, one line each".
Five agents where two would do. Every agent adds calls, latency and cost, and agents that review other agents' reviews add noise, not quality. Start with two. Add a third only when you can name the failure it fixes.
Forgetting that agents share no memory by default. Task outputs flow forward, but an agent does not remember a previous kickoff. Persistent memory across runs exists in CrewAI, but it is a feature you enable deliberately, not something to assume.
Try it yourself
Add a third task, "rewrite the plan as a checklist", owned by a new formatter agent, and extend the scripted replies with a third string. Watch it appear in tasks_output. Then change expected_output on the first task to demand five bullet points and see how the prompt CrewAI builds changes — add verbose=True to the Crew to watch the prompts fly past.
What to learn next
- AutoGen — conversational multi-agent systems — agents that talk in a group chat instead of a task line.
- LangGraph — agents as graphs — the explicit-control-flow alternative.
- Prompt engineering — the skill hiding inside every role and backstory field.
Researcher — Mathematics and papers.
The lineage: role conditioning and society-of-agents
CrewAI is an engineering package around two research threads.
Role conditioning. Assigning a persona measurably changes model behaviour — the mechanism behind role/backstory. The systematic study of cooperative role-played agents is CAMEL (Li et al., 2023, CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, arxiv.org/abs/2303.17760), which paired an "AI user" and "AI assistant" with role prompts and studied drift, termination failure and instruction flipping — all failure modes you will still meet in a crew today.
Structured multi-agent workflows. MetaGPT (Hong et al., 2023, arxiv.org/abs/2308.00352) encoded human standard operating procedures — product manager writes a spec, engineer implements, reviewer checks — as an agent assembly line, showing that constraining who speaks when, with what artefact, beats free-form chat between the same models. CrewAI's task pipeline is this insight productised. Generative Agents (Park et al., 2023, arxiv.org/abs/2304.03442) supplied the memory-and-reflection framing that inspired agent "memory" features generally.
Does adding agents actually help?
The evidence is mixed and task-dependent, and you should hold both results:
- Multi-agent debate improves factuality and reasoning on some benchmarks (Du et al., 2023, Improving Factuality and Reasoning in Language Models through Multiagent Debate, arxiv.org/abs/2305.14325).
- Much of the gain is explainable as extra samples, not extra sociology: majority voting over repeated calls of one agent captures a large share of it (Li et al., 2024, More Agents Is All You Need, arxiv.org/abs/2402.05120).
- Failure-mode taxonomies of multi-agent systems (e.g. Cemri et al., 2025, Why Do Multi-Agent LLM Systems Fail?, arxiv.org/abs/2503.13657) attribute most failures to specification and coordination — poorly defined roles and hand-offs — rather than to model capability. In CrewAI terms: your
expected_outputstrings are the load-bearing wall.
Cost model
For a sequential crew of $n$ tasks where task $i$ consumes context of length $c_i$ and prior outputs are appended, token cost grows roughly as $\sum_{i=1}^{n} (c_i + \sum_{j<i} o_j)$. Where $o_j$ is the output length of task $j$ — earlier outputs are re-read by every later task, so cost is superlinear in pipeline depth. Hierarchical process inserts manager calls between tasks, adding $O(n)$ further calls. This is why "start with two agents" is an economic statement, not a stylistic one.
Positioning, as of 2026
CrewAI's crews occupy the declarative, opinionated end of the spectrum: fast to stand up, strong conventions, less control over exact control flow. Its Flows feature moves toward event-driven, deterministic orchestration for production paths — a partial convergence with the graph philosophy of LangGraph, which starts from explicit control flow and adds convenience upward. Conversation-centred designs such as AutoGen put message exchange, not a task list, at the centre. The frameworks are converging on the same checklist — roles, tools, memory, human gates, observability — and differ mainly in which abstraction sits at the top.
What to learn next
- AutoGen — conversational multi-agent systems — agents that talk in a group chat instead of a task line.
- LangGraph — agents as graphs — the explicit-control-flow alternative.
- Prompt engineering — the skill hiding inside every role and backstory field.