Skip to content
Back to student guides
CrewAILLMsFrameworks & agents3 levels94 sectionsCovers CrewAI 1.15

The Complete CrewAI Guide

Orchestrate role-based teams of AI agents that plan, delegate and collaborate on tasks. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
14sections
27examples

This is part one of three. It covers everything you need to build a working multi-agent application with CrewAI from an empty folder, not a teaser. By the end you can install the toolchain, scaffold a project, write agents and tasks that produce something you would actually send to a colleague, give an agent a tool, read the errors CrewAI throws at you, and explain the difference between a Crew and a Flow without hand-waving. Mid-level and Senior take the same topics further; nothing here is wasted.

Each section ends with a Try it task. Do them as you go. Multi-agent systems are unusually easy to misunderstand from reading, because the interesting failures are behavioural: an agent loops, an agent invents a tool that does not exist, an agent answers the wrong question confidently. You only learn to recognise those by watching your own run scroll past with verbose=True.

1.15.23version this guide verifies against
3.10 – 3.13supported Python (3.14 is not)
4 nounsAgent, Task, Crew, Process
JSONCthe default project scaffold since 1.15.0
Most of what you will read elsewhere is out of date CrewAI changed substantially between the 0.x releases of 2024–2025 and the 1.x line, and 1.0.0 only shipped in October 2025. Blog posts, course videos and the answers a chat model gives you from memory overwhelmingly describe 0.x: the YAML-only project layout, memory split into four kinds, tool-result caching on by default, a code-interpreter tool that no longer exists. This guide is written against 1.15.23, released 28 September 2026, and says so wherever the shipped behaviour differs from what you have probably seen.

What CrewAI is, and the problem it solves

CrewAI is a Python framework for building applications in which several large-language-model agents work together on a job that is too big, too multi-step or too varied for a single prompt. You describe each agent by who it is and what it is for, you describe the work as a list of tasks, and the framework runs the agents, passes one task's output into the next task's context, lets them call tools, and hands you a result object at the end.

The problem it solves is the gap between a prompt and a program. A single call to a language model is a function: text in, text out. That is enough for summarising a paragraph. It is not enough for "research this market, check the numbers against a second source, and write me a two-page brief with citations", because that job has stages, each stage needs different instructions, some stages need to fetch data from the outside world, and the output of each stage has to be checked before it feeds the next.

Before frameworks like this, you wrote that orchestration by hand. The code was not hard, exactly, but it was tedious and it was all the same every time: a loop that calls the model, parses out whether the model wants to use a tool, runs the tool, appends the result to the message history, calls the model again, and gives up after N rounds so a confused model cannot spin forever. On top of that you built your own prompt templates, your own retry logic, your own way of keeping the second stage's prompt from accidentally dropping the first stage's findings, and your own logging so you could tell what went wrong.

CrewAI's bet is that this scaffolding should be a library, and that the natural way to describe the work is by role. You do not write "call the model with this system prompt, then that one". You write: here is a Researcher whose goal is to find verifiable facts about a topic; here is an Analyst whose goal is to turn findings into a brief; here are the two tasks; run them in order. The framework turns that into prompts, runs the tool loop for each agent, and threads the outputs together.

That framing is why it reads so easily and also why it bites beginners. Describing agents in natural language feels like configuration, so people write three vague sentences and expect a system. In practice the role, goal and backstory text is the prompt, and a vague prompt gives a vague system. A large part of becoming competent with CrewAI is learning to write those fields as specifications rather than as flavour.

AGENTwho does it
→
TASKwhat to do
→
CREWwho, in what order
→
CREWOUTPUTwhat you got

Two things worth noticing before you write a line of code. First, CrewAI runs in your process: it is an ordinary Python library, not a server you deploy and talk to, and there is no scheduler or job queue underneath. When crew.kickoff() returns, everything has already happened on your machine. Second, every step an agent takes is a paid LLM call. A crew with three agents, four tasks and a tool each can easily make twenty model calls in one run. Keeping that number visible is a skill, and we come back to it.

Try it
  1. Pick a piece of knowledge work you do by hand that takes more than one step — a weekly report, a candidate screen, a competitor check.
  2. Write it out as stages, one line each, and mark what each stage needs that is not already in your head (a web search, a spreadsheet, a document).
  3. Give each stage a job title.
a list of three to five stages with titles. That list is the shape of a crew, and you will build it for real by the end of this guide.

The four nouns

Almost everything in CrewAI is built from four objects. Learn these precisely and the rest of the API is discoverable.

An Agent is an autonomous worker driven by a language model. It has three required fields — role, goal and backstory — and they are all prompt text. Optionally it gets an llm (which model it thinks with), tools (what it can do besides think), and limits. Internally the agent runs a loop called the executor: the model is asked what to do, and if it asks for a tool the tool is run and the result is fed back, repeating until the model produces a final answer or the loop hits max_iter. In CrewAI 1.15.23 that default is 25 iterations in the shipped code, even though the Agents documentation page still says 20; set it explicitly and you never have to care who is right.

A Task is a unit of work. It has two required fields, description (what to do) and expected_output (what a finished answer looks like), and it is normally assigned to an agent. Beyond that it can declare context — a list of other Task objects whose outputs should be made available to this one — plus its own tools, a structured output type, validation via guardrail, and an output_file to write the result to disk. The expected_output field is the single most under-used lever for beginners: it is what the model is told to aim at, so "a 300-word brief in markdown with a bulleted findings list and at least two source URLs" produces something usable where "a report" produces a shrug.

A Crew is a team: a list of agents, a list of tasks, and a Process. You run it with crew.kickoff(inputs={...}). The inputs dictionary does two jobs at once — it fills {placeholder} templates anywhere in your agent and task text, and it is the only sanctioned way to get runtime values into the run. What comes back is a CrewOutput object, not a string: result.raw is the text, result.pydantic and result.json_dict hold structured output when you asked for it, result.tasks_output is the list of per-task results, and result.token_usage is what the run cost in tokens.

A Process is the execution strategy, and there are exactly two. Process.sequential is the default and the one to learn first: tasks run in the order you listed them, and each task's output automatically becomes context for the next. Process.hierarchical instead puts a manager in charge — you supply either manager_llm or manager_agent — and the manager plans, delegates each task to whichever agent it judges suitable, and validates the result before moving on. There is no third option; if you read about a "consensual" process, you are reading about a plan that was never shipped.

Agent Task Crew Process
Answers Who does the work What the work is Who works together, in what order How the order is decided
Required role, goal, backstory description, expected_output agents, tasks —
Produces A final answer, after a tool loop A TaskOutput A CrewOutput —
Beginner default one model, no tools one agent, clear expected output Process.sequential sequential

Two more nouns appear often enough that you should recognise them now even though you will not configure them today. Tools are callables you hand an agent so it can search the web, read a file or call your API — the LLM decides when to invoke them and with what arguments. Knowledge is a set of documents attached to an agent or crew and queried automatically, which is CrewAI's built-in retrieval-augmented generation; Memory is the separate, unified store of things the system chooses to remember between runs.

Try it
  1. Take the stage list from the previous section and write, for one stage, a one-sentence role, a one-sentence goal, and a two-sentence backstory.
  2. Then write its expected_output so specifically that two different people would produce nearly the same artefact from it.
  3. Count the words in your expected_output. Under fifteen is usually too vague.
a first pair of agent and task definitions in plain text. You will paste them into a real project shortly.

Crews and Flows: the second half of the model

Crews give you autonomy: you say what the goal is, and the agents decide how many tool calls and reasoning steps to take. That is exactly what you want for research and drafting, and exactly what you do not want for the parts of your application that must behave the same way every time, such as deciding which branch of your pipeline to run, or saving a record.

For that, CrewAI has a second primitive: a Flow. A Flow is a Python class that inherits from Flow[StateModel] and whose methods are wired together with decorators: @start() marks an entry point, @listen(...) runs a method after another finishes, and @router(...) returns a string that decides which listener fires next. Combinators or_ and and_ let a step wait for either or both of two predecessors. A Flow carries state — either a plain dictionary or a Pydantic model, always with an automatic UUID id — so steps share data by reading and writing self.state rather than by passing arguments. The flow's result is the return value of whichever method completes last.

The official production guidance is worth internalising early, because it saves you from the most common architectural mistake: use a Flow as the backbone and call Crews, or single agents, from inside Flow steps. The deterministic skeleton — routing, validation, persistence, waiting for a human to approve something — lives in the Flow, where you can test it. The open-ended parts live in Crews, where non-determinism is the point. Beginners who build everything as one large hierarchical crew end up with a system whose control flow is decided by a language model, which is impossible to debug.

flow_example.py
from pydantic import BaseModel
from crewai.flow import Flow, start, listen, router, or_


class State(BaseModel):
    topic: str = ""
    report: str = ""


class BriefFlow(Flow[State]):
    @start()
    def begin(self):
        self.state.topic = self.state.topic.strip()

    @router(begin)
    def pick_depth(self):
        return "long" if len(self.state.topic) > 10 else "short"

    @listen("long")
    def deep_pass(self):
        self.state.report = f"Detailed brief on {self.state.topic}"

    @listen(or_("short", deep_pass))
    def finish(self):
        return self.state.report or f"Short note on {self.state.topic}"


print(BriefFlow().kickoff(inputs={"topic": "AI agents in logistics"}))

Three details make Flows more than syntax. A flow can be persisted with the @persist decorator, which stores its state (SQLite by default) so a run can be resumed later with kickoff(inputs={"id": "<uuid>"}). It can pause for a human using the @human_feedback decorator, which is how you build an approval step without inventing your own queue. And flow.plot("my_flow") writes an interactive HTML diagram of the wiring, which is the fastest way to confirm that the graph in your head matches the graph in your code.

A third, independent axis is how you write a crew down. Since version 1.15.0 the default scaffold is declarative: the crew is described in JSONC files (JSON with comments) and loaded with crewai.project.load_crew. The older classic layout — a Python class decorated with @CrewBase plus config/agents.yaml and config/tasks.yaml — is still fully supported behind the --classic flag. Both produce the same objects. Start with the JSONC default, because that is what crewai create crew gives you and what the current documentation walks through.

A rule of thumb for choosing If you can draw the steps as a diagram with arrows and conditions, that belongs in a Flow. If the honest answer to "how many steps will this take?" is "depends what it finds", that belongs in a Crew. Most real applications are a small Flow calling one or two small Crews.
Try it
  1. Save the flow above as flow_example.py and run it with python flow_example.py. It makes no LLM calls, so it needs no API key.
  2. Change the topic to a short word and run it again. Watch the router send you down the other branch.
  3. Add BriefFlow().plot("brief") and open the generated HTML.
two different printed results from the same code, and a diagram with a fork in it. You have now used the deterministic half of CrewAI before spending a cent on tokens.

Installing CrewAI and checking the setup

You need Python 3.10, 3.11, 3.12 or 3.13. Python 3.14 is explicitly not supported at 1.15.23, which is a real trap if your machine just upgraded itself.

BASH
python3 --version     # macOS and Linux
py --version          # Windows

The officially supported path uses uv, Astral's Python package manager. This is not a stylistic preference: CrewAI's own CLI shells out to uv for crewai install and crewai run, so a project set up another way will fight you.

BASH
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows PowerShell
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Now install the CrewAI command-line tool globally, in its own isolated environment, so it never collides with a project's dependencies.

BASH
uv tool install crewai
uv tool update-shell        # only if you get a PATH warning; then open a new shell

Verify two things, because they are two different installations and they drift apart:

BASH
uv tool list       # shows: crewai v1.15.23  (plus its executables)
crewai version     # the CLI's own version; add --tools to also see crewai-tools
BASH
# inside a project, the library version is separate from the CLI version
uv pip show crewai

That distinction is the single most confusing thing about CrewAI's setup, so state it to yourself clearly: the global CLI is what provides the crewai command, and each project has its own .venv with its own copy of the library. Upgrading one does not upgrade the other, and crewai install syncs a project to its lockfile without ever bumping a version. Keep both on the same release, because the crewai, crewai-core, crewai-cli and crewai-tools packages are version-locked to each other.

BASH
uv tool install crewai --upgrade        # the global CLI
uv add "crewai[tools]>=1.15.23"         # the project: edits pyproject.toml and uv.lock
crewai install                          # sync the project venv

If you only want the library — you are adding agents to an existing application and do not need the scaffolding — skip the tool install entirely:

BASH
uv add crewai                        # or: pip install crewai
uv add "crewai[tools]"               # the built-in tool collection
uv add "crewai[anthropic]"           # native Anthropic SDK
uv add "crewai[google-genai]"        # native Gemini
uv add "crewai[litellm]"             # anything with no native provider

Those extras matter more than they look. Since 1.12.0 CrewAI talks to a set of providers natively — OpenAI, Anthropic/Claude, Azure, Google/Gemini, Bedrock/AWS, OpenRouter, DeepSeek, Ollama, vLLM via hosted_vllm, Cerebras, DashScope and Snowflake — and anything outside that list needs the optional LiteLLM fallback. OpenAI itself is a core dependency and needs no extra.

Windows: the C++ build tools If installation fails while compiling a dependency (historically chroma-hnswlib), install Visual Studio Build Tools with the Desktop development with C++ workload and retry. Also run uv tool update-shell and restart the terminal, or the crewai command will not be found.

Finally, the version numbers you will see in old material. The installation page's sample output still shows crewai v0.102.0, and a note there says CrewAI 0.175.0 requires openai >= 1.13.3. Both are stale: 1.15.23 requires openai>=2.30.0. Treat the version you can see with crewai version as the truth.

Try it
  1. Install uv, then uv tool install crewai.
  2. Run crewai version --tools and write the two numbers down.
  3. Run python3 -c "import crewai" outside a project and watch it fail — the tool install is deliberately isolated.
a working crewai command and a clear sense that the CLI and the library are two separate things. That understanding prevents a whole family of "but I upgraded it" problems.

Your first crew, step by step

Scaffold a project. The command prompts you for a provider and a model, then writes the API key you give it into .env.

BASH
crewai create crew research_crew
cd research_crew
crewai install

Use --provider <name> to preselect a provider non-interactively, or --skip-provider to skip the prompt entirely and fill .env in yourself. If you want the older Python-and-YAML layout, add --classic.

What you get is a JSON-first project:

TEXT
research_crew/
├── .env
├── .gitignore
├── agents/
│   └── researcher.jsonc
├── crew.jsonc
├── knowledge/
├── skills/
├── tools/
├── pyproject.toml
└── README.md

An agent lives in its own file under agents/, and the file name is the agent's name everywhere else:

agents/researcher.jsonc
{
  "role": "Senior Research Specialist for {topic}",
  "goal": "Find verifiable, recent facts about {topic} and keep the source for each one",
  "backstory": "A careful analyst who distrusts single sources and always notes the date of a claim.",
  "llm": "openai/gpt-4.1-mini",
  "settings": { "verbose": true, "allow_delegation": false }
}

The crew file names the agents, lists the tasks in order, and sets the process:

crew.jsonc
{
  "name": "Research Crew",
  "agents": ["researcher", "analyst"],
  "tasks": [
    {
      "name": "research_task",
      "description": "Research the current state of {topic}. Prefer sources from the last 18 months.",
      "expected_output": "Eight findings. Each one sentence, each followed by its source URL.",
      "agent": "researcher"
    },
    {
      "name": "analysis_task",
      "description": "Turn the findings into a brief for a non-specialist decision maker.",
      "expected_output": "A 400-word markdown brief: a two-sentence summary, four themes, and the source list.",
      "agent": "analyst",
      "context": ["research_task"],
      "output_file": "output/report.md",
      "markdown": true
    }
  ],
  "process": "sequential",
  "verbose": true,
  "inputs": { "topic": "Artificial Intelligence in Healthcare" }
}

Three rules govern these files and each one is a real error if you break it. Every name in agents must match a file in agents/. Every task's agent must match one of those names. And every {placeholder} in any agent or task text must either appear in inputs or be supplied at run time — crewai run will prompt you for anything missing rather than silently leaving the literal braces in the prompt.

Then run it:

BASH
crewai run

With verbose on you will see each agent announce its role, think, optionally call a tool, and produce a final answer, then the next task start with the previous output in its context. A fresh scaffold writes output/report.md when it finishes. Open that file: the point of output_file is that the deliverable lands on disk rather than scrolling off your terminal.

The same crew can be loaded from Python when you want to embed it in an application:

run_crew.py
from pathlib import Path
from crewai.project import load_crew

crew, default_inputs = load_crew(Path("crew.jsonc"))
result = crew.kickoff(inputs={**default_inputs, "topic": "AI in radiology"})
print(result.raw)
print(result.token_usage)
Read the terminal, not just the result The verbose log is the only honest account of what happened. Count the agent steps and the tool calls in your first run. That number, multiplied by your per-call cost, is your application's economics, and it is much easier to control on day one than after you have built three layers on top.
Try it
  1. Scaffold research_crew, add a second agent file agents/analyst.jsonc, and paste the crew.jsonc above.
  2. Run crewai run and read output/report.md.
  3. Now delete "topic" from inputs and run again.
a finished report, and then a prompt asking you for the topic. You have seen both halves of how inputs reach a crew.

The same crew in plain Python

The JSONC scaffold is convenient, but you should be able to write a crew in a single file, because that is how you will experiment and how most examples you read are written.

minimal_crew.py
from crewai import Agent, Task, Crew, Process, LLM

llm = LLM(model="openai/gpt-4.1-mini", temperature=0.2)

researcher = Agent(
    role="Researcher",
    goal="Find verifiable facts about {topic}",
    backstory="A careful analyst who always records the source of a claim.",
    llm=llm,
    max_iter=25,
    verbose=True,
)

research = Task(
    description="Research {topic} and note the date of each claim.",
    expected_output="Five findings, one sentence each, each with a source URL.",
    agent=researcher,
)

crew = Crew(agents=[researcher], tasks=[research], process=Process.sequential, verbose=True)

result = crew.kickoff(inputs={"topic": "AI agents"})
print(result.raw)
print(result.token_usage)

Note the model string. CrewAI always wants the provider/model-id form — openai/gpt-4.1-mini, anthropic/claude-sonnet-4-6, ollama/llama3 — because the prefix is how it decides which provider implementation to use. A bare model name is the most common reason a beginner's first script cannot find a provider at all.

If you set no llm, CrewAI resolves one from the environment in order: MODEL, then MODEL_NAME, then OPENAI_MODEL_NAME, and failing all three it falls back to gpt-4.1-mini. That fallback is why a script with no OpenAI key can fail with an OpenAI authentication error even when you thought you had configured Anthropic everywhere. (The documentation table that says the default is gpt-4 is stale.)

For jobs that need no team at all, an Agent can be used alone. This is the cheapest way to test whether your role and goal text is any good, because there is no task plumbing in the way:

solo_agent.py
from crewai import Agent

agent = Agent(
    role="Editor",
    goal="Rewrite text so a busy executive can read it in thirty seconds",
    backstory="Twenty years of cutting other people's paragraphs in half.",
)

out = agent.kickoff("Rewrite: our Q3 initiative achieved synergies across verticals.")
print(out.raw)

agent.kickoff(...) returns a LiteAgentOutput with raw, pydantic, agent_role and usage_metrics — not a CrewOutput. There is an async twin, agent.kickoff_async(...), and the crew equivalents are crew.akickoff(...) for native async and crew.kickoff_async(...) for a thread-based version. You need one of them if you are calling CrewAI from inside FastAPI or a Jupyter notebook, because invoking the synchronous kickoff from within a running event loop raises an error that tells you exactly this.

To run the same crew over many inputs, crew.kickoff_for_each(inputs=[{...}, {...}]) loops for you and returns a list of outputs, with kickoff_for_each_async and akickoff_for_each as the concurrent versions.

Try it
  1. Run minimal_crew.py with a real key in your environment.
  2. Print result.tasks_output[0].raw and compare it with result.raw.
  3. Delete the llm= argument, unset MODEL, and run again.
identical text in both prints for a one-task crew, and then a run that quietly switches to gpt-4.1-mini. Knowing where the default comes from is worth the experiment.

Writing agents and tasks that actually work

Everything about the quality of a crew's output traces back to four text fields. Here is what separates a definition that works from one that produces mush.

Specific

  • role: "Clinical-trials research specialist"
  • goal: "Find trials registered in the last 18 months and record phase, sponsor and status"
  • backstory: states what the agent distrusts and what it refuses to guess
  • expected_output: "A markdown table with columns trial ID, phase, sponsor, status, source URL"

Vague

  • role: "Researcher"
  • goal: "Research the topic well"
  • backstory: "You are an expert with many years of experience."
  • expected_output: "A detailed report"

The goal should name the decision the output feeds. The backstory is where you put constraints and taste — "never state a number you cannot attribute", "prefer primary sources", "write for a reader who is not a doctor" — because that text is in the system prompt for every step the agent takes. The expected_output is the shape of the artefact, and it is the field people leave weakest even though it has the most direct effect on whether the result is usable.

Beyond text, a handful of Task features turn a demo into something you can rely on.

Context makes dependencies explicit. In a sequential crew each output flows to the next task automatically, but context=[research_task, pricing_task] says precisely which earlier outputs this task should see. Order matters: a task may not reference a task that comes later, and CrewAI refuses the run with a message naming the forward dependency rather than silently producing nonsense.

Structured output replaces parsing prose. Pass a Pydantic model as output_pydantic and the finished result is available as result.pydantic:

structured.py
from pydantic import BaseModel
from crewai import Agent, Task, Crew


class Brief(BaseModel):
    summary: str
    themes: list[str]
    sources: list[str]


analyst = Agent(role="Analyst", goal="Summarise findings", backstory="Blunt and brief.")

task = Task(
    description="Summarise the findings about {topic}.",
    expected_output="A summary, three themes, and the source URLs.",
    agent=analyst,
    output_pydantic=Brief,
)

result = Crew(agents=[analyst], tasks=[task]).kickoff(inputs={"topic": "AI agents"})
print(result.pydantic.themes)

Two rules here, both of which CrewAI enforces with a clear error. Pass the class (Brief), never an instance (Brief(...)). And set either output_pydantic or output_json, never both.

Guardrails validate the output and retry if it fails. A guardrail is either a function taking one TaskOutput and returning a (bool, value) tuple, or a plain string that becomes an LLM-judged check using the task's own agent. Failures retry up to guardrail_max_retries, which defaults to 3.

guardrail.py
from typing import Any
from crewai import TaskOutput


def has_sources(result: TaskOutput) -> tuple[bool, Any]:
    if "http" not in result.raw:
        return (False, "Output contained no source URLs. Add at least two.")
    return (True, result.raw)


# Task(..., guardrail=has_sources)
# Task(..., guardrail="Must cite at least two URLs")
# Task(..., guardrails=[has_sources, "Under 400 words"])

Human input is the simplest safety valve there is: Task(..., human_input=True) pauses and asks you to review the output before the crew moves on. For an application rather than a terminal session, the Flow decorator @human_feedback is the grown-up version.

output_file writes the result to disk, creating the directory when you pass create_directory=True. CrewAI validates this path strictly — no .., no ~, no shell characters — and refuses anything that looks like path traversal, so use plain relative paths like output/report.md.

Try it
  1. Rewrite your weakest expected_output so it names a format, a length and a required element.
  2. Add the has_sources guardrail to that task and run the crew.
  3. Deliberately phrase the description so sources are unlikely, and watch the retry.
a visible retry in the log with your own failure message fed back to the agent. That loop is CrewAI's main quality mechanism, and you now control it.

Giving an agent tools

An agent with no tools can only reason over what is already in its prompt. Tools are how it reaches the world: search the web, read a file, call your internal API. The agent's model decides when to call a tool and what arguments to pass, which is both the power and the risk.

The built-in collection ships in crewai-tools, installed with the tools extra. It includes SerperDevTool for web search (needs SERPER_API_KEY), ScrapeWebsiteTool, FileReadTool, FileWriterTool, DirectoryReadTool, PDFSearchTool, CSVSearchTool, WebsiteSearchTool, ExaSearchTool and GithubSearchTool, among others.

BASH
uv add "crewai[tools]"
with_tools.py
from crewai import Agent, Task, Crew
from crewai_tools import SerperDevTool

search = SerperDevTool()

researcher = Agent(
    role="Market researcher",
    goal="Find and cite current facts about {topic}",
    backstory="Never states a figure without a source URL.",
    tools=[search],
    verbose=True,
)

task = Task(
    description="Search for recent developments in {topic}.",
    expected_output="Six findings, each with a source URL.",
    agent=researcher,
)

print(Crew(agents=[researcher], tasks=[task]).kickoff(inputs={"topic": "agentic AI"}).raw)

Writing your own tool is a small class. Note the import path: it is from crewai.tools import BaseTool, tool. The old crewai_tools.BaseTool and crewai.agents.tools.tool paths no longer exist, and a surprising amount of published example code still uses them.

tools/currency.py
from crewai.tools import BaseTool
from pydantic import BaseModel, Field


class ConvertInput(BaseModel):
    amount: float = Field(..., description="Amount in US dollars")
    rate: float = Field(..., description="Dollars per unit of target currency")


class ConvertTool(BaseTool):
    name: str = "Convert USD"
    description: str = "Convert an amount in US dollars to another currency at a given rate."
    args_schema: type[BaseModel] = ConvertInput

    def _run(self, amount: float, rate: float) -> str:
        return f"{amount / rate:.2f}"

For something trivial, the decorator form is shorter, and the docstring becomes the description:

quick_tool.py
from crewai.tools import tool


@tool("Word count")
def word_count(text: str) -> str:
    """Count the words in a piece of text and return the number."""
    return str(len(text.split()))

The description and the field descriptions are not documentation; they are the only thing the model reads when deciding whether this tool is the right one. A vague description is the direct cause of the two most common tool failures: the agent calls the wrong tool, or it invents a tool name that does not exist and CrewAI replies with You tried to use the tool {tool}, but it doesn't exist.

Three current facts about tools that older material gets wrong. Tool-result caching is off by default: Crew.cache has defaulted to False since 1.14.3, and Agent.cache=True only means that agent participates when caching is enabled on the crew. Enable Crew(cache=True) deliberately, and only for tools whose results do not change. There is no built-in code execution any more: CodeInterpreterTool was removed in 1.14.0, and Agent.allow_code_execution and code_execution_mode are deprecated no-ops whose warning points you to external sandboxes such as E2B or Modal. And MCP is built in: Model Context Protocol servers are attached with Agent(mcps=[...]), taking either string references or MCPServerStdio, MCPServerHTTP and MCPServerSSE objects from crewai.mcp. If you are curious about that protocol, the MCP guide covers it from first principles.

Least privilege is not optional The model chooses the arguments. A file-writing tool with access to your whole home directory is a file-writing tool pointed at your whole home directory by a system that can be talked into things by a web page it read. Give each agent the narrowest tools that let it finish its task, and treat anything an agent fetched from the internet as untrusted text rather than as instructions.
Try it
  1. Add the word_count tool to an agent and write a task that forces it to be used.
  2. Now rewrite the docstring to something useless, like "Does a thing", and run again.
  3. Compare the two verbose logs.
with a clear description the tool is called immediately; with a bad one the agent guesses, or skips it. Tool descriptions are prompts.

The everyday commands, grouped by what you want

You will use perhaps eight commands regularly. Grouping them by intent is easier than memorising a list.

Starting something. crewai create crew NAME scaffolds a JSON-first crew, with --classic for the YAML layout, --provider to preselect a provider and --skip-provider to skip the prompt. crewai create flow NAME scaffolds a Python Flow project, with --declarative for the JSON-described variant. crewai create tool NAME and crewai create skill NAME add a custom tool or a reusable instruction package.

Running it. crewai install is uv sync into the project's .venv. crewai run reads [tool.crewai] in pyproject.toml to decide whether this project is a crew or a flow and runs it; for flows it accepts --definition PATH and --inputs '{"topic":"AI"}'. crewai chat opens an interactive session with your crew, which requires a chat_llm set on the crew.

Looking at state. crewai version, with --tools to include crewai-tools. crewai flow plot draws the flow graph. crewai log-tasks-outputs lists the task IDs from the most recent kickoff, and crewai replay -t TASK_ID re-runs from that task onward — the fastest way to iterate on a late task without paying for the early ones again.

Clearing state. crewai reset-memories needs at least one flag or it refuses: -m for memory, -kn for knowledge, -akn for agent knowledge, -k for the latest kickoff outputs, -a for everything.

BASH
crewai reset-memories -m        # the unified memory store
crewai reset-memories -a        # everything, after changing an embedder
The old reset flags are deprecated Memory was unified into a single store around 1.10. The old -l (long), -s (short) and -e (entity) flags still exist but are hidden and deprecated: they print Warning: --long is deprecated. Use --memory (-m) instead. All memory is now unified. Any tutorial using them predates the change, which is a useful dating signal for the rest of that tutorial too.

Two more command families exist, and you should know they are there even though a beginner does not need them. crewai train -n 5 runs iterations collecting your feedback into trained_agents_data.pkl, which agents then load as suggestions, and crewai test -n 3 runs a crew repeatedly and scores it with a model. And a set of commands — crewai login, crewai deploy create, crewai deploy push, crewai deploy logs, crewai traces enable — connect to CrewAI AMP, the commercial Agent Management Platform at app.crewai.com that offers hosted deployment, a REST endpoint per crew, and trace viewing. None of it is required to use the framework.

Also note the CLI's recent renames, because the old spellings appear everywhere: flags became kebab-case (--n-iterations, --task-id, --skip-provider), crewai tool create became crewai create tool, and crewai flow kickoff became crewai run — the old form still works but prints a deprecation notice.

Try it
  1. Run crewai log-tasks-outputs after a multi-task run.
  2. Pick the last task's ID and run crewai replay -t <id>.
  3. Time both runs.
a replay that is dramatically faster and cheaper than a full kickoff. This is the loop you will live in while tuning a late-stage prompt.

Configuration: models, keys, memory and knowledge

Configuration is mostly environment variables plus a handful of constructor arguments, and the thing to understand is not the list but where hidden model calls come from.

Your .env needs a provider key and, usually, a model:

.env
OPENAI_API_KEY=sk-...
MODEL=openai/gpt-4.1-mini
SERPER_API_KEY=...

Other providers follow the same shape: ANTHROPIC_API_KEY, GEMINI_API_KEY or GOOGLE_API_KEY, AZURE_API_KEY with AZURE_ENDPOINT, AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY with a region, SNOWFLAKE_PAT with SNOWFLAKE_ACCOUNT_URL. Always write the model as provider/model-id. The scaffold's .gitignore excludes .env; keep it that way.

Now the part that costs people money. Several CrewAI features default to an OpenAI model independently of the agent you configured. Crew planning (planning_llm), the memory system's analysis model, and the default for crewai test -m all default to gpt-5.4-mini in the 1.15.23 code, and memory's default embedder is OpenAI's text-embedding-3-large while knowledge's is text-embedding-3-small. So a crew where every agent runs on Anthropic can still fail without OPENAI_API_KEY, or quietly bill OpenAI. If you are standardising on one provider, override every one of them: each agent's llm, plus planning_llm, manager_llm, Memory(llm=..., embedder=...), and the knowledge embedder.

crew_config.py
from crewai import Crew, Process

crew = Crew(
    agents=[...],
    tasks=[...],
    process=Process.sequential,
    verbose=True,
    memory=False,                  # default; True needs an embedder and a key
    cache=False,                   # default since 1.14.3
    max_rpm=20,                    # rate-limit the whole crew
    output_log_file=True,          # writes logs.txt
)

The Crew defaults worth knowing by heart at this level: process=sequential, verbose=False, memory=False, cache=False, planning=False, share_crew=False. On the Agent: max_iter=25, allow_delegation=False, max_rpm=None, max_execution_time=None, max_retry_limit=2, respect_context_window=True (which auto-summarises when a prompt would overflow the model's window, at the cost of extra calls). Note that verbose is now a plain boolean — passing an integer to select a log level is a 0.x habit that no longer does anything.

Knowledge is the feature to reach for when your agents need to know about your documents. Attach sources at crew or agent level and CrewAI embeds them into ChromaDB and queries them automatically:

with_knowledge.py
from crewai import Crew
from crewai.knowledge.source.string_knowledge_source import StringKnowledgeSource

policy = StringKnowledgeSource(content="Refunds are issued within 14 days of purchase.")

crew = Crew(agents=[...], tasks=[...], knowledge_sources=[policy])

File-based sources exist for text, PDF, CSV, Excel and JSON, and their paths are resolved relative to the project's knowledge/ directory, which is the single most common cause of a "file not found" here. The store itself lives in a per-OS application-data directory — ~/Library/Application Support/CrewAI/{project}/knowledge/ on macOS, ~/.local/share/CrewAI/{project}/knowledge/ on Linux, %LOCALAPPDATA%\CrewAI\{project}\knowledge\ on Windows — and you can move all CrewAI storage by setting CREWAI_STORAGE_DIR.

Memory is different from knowledge and the two are constantly confused. Knowledge is documents you provide, read-only, for retrieval. Memory is what the system records as it runs so a later run can recall it. Since the unification there is one Memory class backed by LanceDB, stored by default in ./.crewai/memory, with memories filed under hierarchical scopes like /project/alpha and recall scored by a blend of semantic similarity, recency and importance. Turning it on is Crew(memory=True); that requires an embedder, which by default means an OpenAI key.

Two operational notes. If you change embedder, reset the stores — old stores were 1536-dimensional, the current default is 3072, and mixing them produces a dimension-mismatch error. And all of this state is local files by default, which matters the moment you put a crew in a container: mount a volume or set CREWAI_STORAGE_DIR, or your memory vanishes between runs.

Data residency, for readers in the Gulf and Egypt Because knowledge and memory are embedded by a hosted provider by default, your documents leave the machine when you enable them. If a client requires data to stay in-region, your options are a regional endpoint from a provider that offers one (Azure OpenAI and Bedrock both have Middle East regions), or a local embedder — Memory(embedder={"provider": "ollama", ...}) keeps embedding on your own hardware. Separately, leave share_crew at its default of False: setting it true sends goals, backstories, inputs and task outputs to CrewAI.
Try it
  1. Add a StringKnowledgeSource containing a fact no model could know, such as an invented internal policy.
  2. Ask your crew a question that depends on it.
  3. Remove the source and ask again.
the right answer first, and a confident invention second. That contrast is the clearest possible demonstration of why retrieval exists.

Reading the errors

CrewAI's error messages are unusually direct, and most of them name the fix. Learning to recognise the common ones turns an afternoon of confusion into a two-minute correction.

Provider and model problems. Unable to initialize LLM with model '<m>'. The model did not match any supported native provider ... and the LiteLLM fallback package is not installed. means your prefix is not one of the native providers; install the fallback with uv add 'crewai[litellm]' or switch providers. A message like Anthropic native provider not available, to install: uv add "crewai[anthropic]" means the right provider but a missing extra. An OpenAI authentication complaint in a crew you thought was all-Anthropic almost always means a hidden default: an agent with no llm, or planning, memory or the test command reaching for OpenAI.

Configuration refusals. Attribute 'manager_llm' or 'manager_agent' is required when using hierarchical process. and Manager agent should not be included in agents list. are the two hierarchical-crew mistakes, in that order. Sequential process error: Agent is missing in the task with the following description: {description} means exactly what it says. Only one output type can be set, either output_pydantic or output_json. and a Pydantic validation error on output_pydantic (you passed an instance instead of the class) are the structured-output pair. Missing required template variable '<var>' in description means a {placeholder} with nothing to fill it.

Ordering refusals. Task '<d>' has a context dependency on a future task '<d2>', which is not allowed. means reorder your tasks. The crew must end with at most one asynchronous task. means make the trailing tasks synchronous or add a synchronous aggregator at the end.

Behavioural failures, which are the interesting ones. If your output ends with the sentence "You can't keep going, here is the best final answer you generated:", the agent hit max_iter. CrewAI makes one final call demanding a best effort and gives you that. The cause is almost never that 25 iterations is too few; it is usually a tool that keeps failing, so the agent keeps retrying it, or a task so broad the agent cannot tell when it is done. Fix the tool or tighten the task before you raise the limit.

Invalid response from LLM call - None or empty. means the provider returned nothing: check the model id for a typo, raise max_tokens, consider content filtering, and look at the verbose log. Context length exceeded. Summarizing content to fit the model context window. Might take a while... is respect_context_window doing its job; the fix is less input, not a bigger window. Task '<d>' execution timed out after <n> seconds. is your own max_execution_time.

Tool failures have their own family: You tried to use the tool {tool}, but it doesn't exist. You must use one of the following tools, use one at time: {tools}. is a hallucinated tool name, usually from weak descriptions or a weak model. Error: the Action Input is not a valid key, value dictionary. means the model produced malformed arguments; simplify the args_schema. And I encountered an error while trying to use the tool. This was the error: {error}. is your tool raising an exception — which it should not do in normal operation; return a ToolFailure instead so the agent can react.

Finally, two environment classics. Agent execution was invoked synchronously from within a running event loop. tells you to use kickoff_async or akickoff, and it happens the first time you call a crew from FastAPI or a notebook. And Error. A valid pyproject.toml file is required. from crewai uv ... simply means you are in the wrong directory.

Try it
  1. Set a task's max_execution_time=5 on a research task and watch the timeout message.
  2. Set max_iter=2 on an agent with a tool and find the forced-final-answer sentence in the output.
  3. Pass output_pydantic=Brief() instead of Brief and read the validation error.
three errors you caused on purpose. Causing them deliberately once is the cheapest way to recognise them later under pressure.

Putting it all together

Here is one small end-to-end project that uses everything above: a Flow as the backbone, a two-agent crew inside it, a guardrail, structured output and a file on disk. It is deliberately small enough to read in one sitting.

brief/main.py
from pathlib import Path
from typing import Any

from pydantic import BaseModel
from crewai import Agent, Crew, LLM, Process, Task, TaskOutput
from crewai.flow import Flow, listen, start


class Brief(BaseModel):
    summary: str
    themes: list[str]
    sources: list[str]


def has_two_sources(result: TaskOutput) -> tuple[bool, Any]:
    if result.raw.count("http") < 2:
        return (False, "Fewer than two source URLs. Add more and cite them inline.")
    return (True, result.raw)


def build_crew() -> Crew:
    llm = LLM(model="openai/gpt-4.1-mini", temperature=0.2)

    researcher = Agent(
        role="Research specialist on {topic}",
        goal="Find recent, attributable facts about {topic}",
        backstory="Never states a figure without a source. Notes the date of every claim.",
        llm=llm,
        max_iter=25,
        verbose=True,
    )
    analyst = Agent(
        role="Analyst writing for a non-specialist decision maker",
        goal="Turn findings about {topic} into a brief someone can act on",
        backstory="Cuts jargon. Refuses to pad. Keeps every claim traceable to a source.",
        llm=llm,
        max_iter=25,
        verbose=True,
    )

    research = Task(
        description="Research the current state of {topic}, preferring the last 18 months.",
        expected_output="Six findings, one sentence each, each followed by its source URL.",
        agent=researcher,
        guardrail=has_two_sources,
    )
    write = Task(
        description="Write a brief from the findings for a reader with no background in {topic}.",
        expected_output="A summary of two sentences, three themes, and the list of sources.",
        agent=analyst,
        context=[research],
        output_pydantic=Brief,
        output_file="output/brief.md",
        create_directory=True,
        markdown=True,
    )

    return Crew(
        agents=[researcher, analyst],
        tasks=[research, write],
        process=Process.sequential,
        verbose=True,
    )


class State(BaseModel):
    topic: str = ""
    brief: Brief | None = None


class BriefFlow(Flow[State]):
    @start()
    def normalise(self):
        self.state.topic = self.state.topic.strip() or "AI agents in healthcare"

    @listen(normalise)
    def research_and_write(self):
        result = build_crew().kickoff(inputs={"topic": self.state.topic})
        self.state.brief = result.pydantic
        print("tokens:", result.token_usage)
        return self.state.brief


if __name__ == "__main__":
    brief = BriefFlow().kickoff(inputs={"topic": "AI agents in radiology"})
    print(brief.summary)
    print(Path("output/brief.md").read_text()[:400])

Walk through what happens when you run it. The flow starts, normalises the topic in state, and calls its single listener. That listener builds a crew and kicks it off with the topic as an input, which interpolates {topic} into both agents' role and goal text and both tasks' descriptions. The researcher runs its executor loop, producing findings; the guardrail inspects the raw output and, if fewer than two URLs appear, returns a failure message that is fed back to the agent for another attempt, up to three times. The analyst then runs with the research output in its context, and because the task declares output_pydantic, CrewAI coerces the answer into a Brief and also writes the markdown to output/brief.md, creating the directory. The crew returns a CrewOutput; the flow stores the parsed model in state, prints the token cost, and returns it as the flow's result.

Everything here is a lever you now understand. Want cheaper? Give the researcher a smaller model via its own llm and leave the analyst on the larger one. Want it auditable? Add output_log_file=True to the crew. Want a human in the loop? Add human_input=True to the writing task, or move the approval into the flow with @human_feedback. Want it resumable? Decorate the flow with @persist and resume with kickoff(inputs={"id": "<uuid>"}). Want to run it over fifty topics? Replace the single kickoff with kickoff_for_each.

Try it
  1. Run the project end to end and read output/brief.md.
  2. Note the token usage, then set the researcher to a cheaper model and compare both the cost and the quality.
  3. Add human_input=True to the writing task and run once more.
a finished brief on disk, a number you can put in a budget, and a run that stopped and asked your permission. That is a small but genuinely complete application.

What you can now do, and what comes next

You can install the CrewAI toolchain the supported way and explain why the global CLI and the project library are separate installs. You can scaffold a JSON-first project, read and edit crew.jsonc and the files under agents/, and write the same crew directly in Python. You know the four nouns and the two processes, and you know that a Flow is the deterministic backbone that should own your control flow while crews own the open-ended parts. You can write role, goal, backstory and expected_output text that specifies rather than gestures. You can give an agent built-in or custom tools, with the correct crewai.tools import path. You can enforce output shape with Pydantic and output quality with guardrails. You know where hidden OpenAI calls come from and how to override each one. And you can read the error messages, including the behavioural ones that have no stack trace.

What you have deliberately not covered: the hierarchical process in anger, async and parallel execution, streaming, event listeners and the event bus, checkpointing and forking runs, crew-level and agent-level planning, declarative flows, MCP servers in depth, and deployment. Those are the mid-level material, and they all build on exactly the model you now have.

Three habits to carry forward. Keep verbose=True until a crew is boring, because the log is the only account of what the agents actually did. Pin your versions and read the changelog, because 1.x ships roughly weekly and defaults have genuinely changed inside the minor line — tool caching, max_iter, the default models. And treat the number of LLM calls per run as a first-class metric from the beginning; it is the difference between a demo and something you can afford to run.

For where to go next in this catalogue: LangGraph is the natural comparison, a graph-first framework where you build the state machine yourself, and reading the two against each other sharpens what CrewAI's role abstraction is actually buying you. Once a crew runs more than a few times a day you will want to see inside it, which is what Langfuse is for — CrewAI's event bus and OpenTelemetry-based tracing integrate with it directly. And when you need to know whether a change made your output better rather than just different, RAGAS covers evaluation. If your crews talk to tools over Model Context Protocol, the MCP guide explains the protocol your mcps=[...] list is speaking.

Sources