This is part one of three, and it is written for someone who has never used smolagents. By the end you will have installed it, pointed it at a language model, built an agent that answers questions by writing and running Python, given that agent your own tools, read its step-by-step trace, and understood why its code executor is a convenience and not a security wall. You will also know the dozen errors that account for most beginner frustration, and which of them are your fault and which are the library's.
Each section ends with a Try it task. Do them as you go. Agents are strange until you have watched one think, stumble and recover on your own screen, and no amount of reading replaces that.
This guide targets smolagents 1.26.0, the current release at the time of writing. The library is still below version 2.0 and removes things in minor releases, so you will see a few places where older tutorials on the internet no longer work. Those places are flagged as they come up.
What smolagents is, and the problem it solves
A language model on its own can only produce text. Ask it today's exchange rate and it will either guess or admit it does not know. An agent is a program that wraps a model in a loop so it can do more than talk: it can decide to search the web, run a calculation, read a file or call your internal API, look at the result, and decide what to do next. The model supplies the judgement; the loop supplies the hands.
Before libraries like this existed, every team wrote that loop by hand. You would prompt the model to answer in a strict format, parse the reply with regular expressions, call the right function, paste the output back into the conversation and repeat until the model said it was finished. It works, but the details are fiddly. What happens when the model answers in the wrong format? When it calls a function that does not exist? When it never stops? Each project rediscovered the same bugs.
smolagents is a small Python library from Hugging Face that packages that loop. The word "smol" is deliberate: the core agent logic is roughly a thousand lines, small enough to read in an afternoon. That matters for a beginner, because when something goes wrong you can open the source and understand it instead of fighting a framework with five layers of abstraction.
Its signature idea is the code agent. Most agent frameworks ask the model to emit actions as JSON, something like a structured request naming a tool and its arguments. smolagents' default agent instead asks the model to write its actions as Python code that calls your tools as ordinary functions. Code is a better action language than JSON for the same reason it is a better way to give instructions to a computer: it can loop, branch, store intermediate results in variables and combine several tool calls in a single step. The project's README cites research showing this approach needs around 30 percent fewer steps, and therefore fewer model calls, than JSON-style tool calling. Treat that as a claim from the project itself, but the intuition is sound. Asking "what is the population of the three largest cities in Egypt, summed?" is one short script for a code agent, where a JSON agent would need several separate round trips.
smolagents also supports the JSON style, through a second agent class, so you can choose. And it is not tied to one model vendor. It works with models served through Hugging Face's Inference Providers, with OpenAI and any server that speaks the OpenAI protocol, with LiteLLM (a translation layer covering over a hundred providers, including a local Ollama), with Azure and Amazon Bedrock, and with models that run on your own machine through Transformers.
- Pick a question you would normally answer by opening a browser, copying a number and doing arithmetic, for example the average of three currency rates.
- Write down every distinct action you would take: each search, each calculation.
- Mark which actions depend on the result of an earlier one.
The mental model: four nouns
Everything in smolagents is built from four ideas. Learn these and the rest of the documentation becomes readable.
The first is the model. This is the language model that does the thinking, wrapped in a small Python class. smolagents ships several wrappers. InferenceClientModel talks to Hugging Face's Inference Providers; OpenAIModel talks to OpenAI or any compatible server; LiteLLMModel talks to anything LiteLLM supports; TransformersModel runs a model on your own hardware. They all behave the same from the agent's point of view: it hands over a list of chat messages and gets one reply back.
The second is the tool. A tool is a Python function the agent is allowed to use, plus a description the model can read. The model never sees your function's code. It sees only the tool's name, a sentence explaining what it does, and the names and types of its inputs. From that, it decides whether and how to call it. This is why tool descriptions matter so much: they are the entire user manual the model gets.
The third is the agent. The agent owns the loop. You construct it with a model and a list of tools, then call agent.run("your task"). The two agent classes you will use are CodeAgent, which has the model write Python, and ToolCallingAgent, which uses the model provider's native JSON tool-calling.
The fourth is memory. As the agent works it records every step: your task, each thing the model said, each piece of code or tool call, each result and each error. At every step the whole record is sent back to the model so it knows what has happened so far. Memory is also what you inspect afterwards to understand what the agent did and why.
One more piece lives inside the code agent and deserves a name now: the executor. When a CodeAgent receives Python from the model, something has to run it. That something is the executor, and by default it is a custom Python interpreter that runs inside your own process with a restricted set of allowed imports. Later in this guide you will see exactly what it allows and where its limits are.
Finally, every agent always has one built-in tool you did not write, called final_answer. When the model calls it, the loop stops and the argument becomes the value returned from run(). If you remember nothing else about control flow: the agent keeps looping until the model calls final_answer or until it runs out of steps.
- Without looking back, write the four nouns and one sentence for each.
- For the task "convert 250 US dollars to Egyptian pounds", say what the tool would be and what the model would do.
How one run works: the loop
Agents in smolagents follow a pattern called ReAct, short for reasoning and acting. At each step the model first writes a short piece of reasoning, then takes an action, then sees what the action produced. Calling agent.run(task) does the following, in this order.
- The task is stored in memory as the first entry.
- If you enabled planning (it is off by default), the agent writes a plan first.
- The agent then repeats an action step until it is finished. In each action step it converts everything in memory into chat messages, calls the model, extracts the action from the reply, executes it, and records what happened.
- For a
CodeAgent, the action is a block of Python code. For aToolCallingAgent, it is one or more tool calls in JSON. - The loop ends when the model calls
final_answer, or when the step limit is reached.
The step limit is controlled by max_steps, which defaults to 20. If the agent reaches it without finishing, it does not simply crash. It makes one extra model call asking for a best-effort answer based on what it has seen, and returns that. This is a guard against runaway loops, and runaway loops cost money because every step is a paid model call.
That last point deserves emphasis for a beginner. Every step sends the full memory to the model again. Step one sends the task. Step five sends the task plus four steps of code and output. The cost and the time of a run therefore grow faster than the number of steps. A task that takes three steps is cheap; one that wanders for twenty is not. Much of the skill of building agents is keeping runs short, and you will see the levers for that in the tips guide.
A code agent's reply has a recognisable shape. It begins with a line of thought, then a block of code between <code> and </code> tags. A typical reply looks like this:
Thought: I need the current population figure, so I will search for it first.
<code>
results = web_search("population of Cairo")
print(results)
</code>
The executor runs that code and captures what it printed. Printing is how the model "looks at" a value: only printed output and the final answer return to the model. So you will often see the agent write print(...) on purpose. The captured text becomes an observation, which goes into memory and into the next model call. On the following step the model might write code that extracts a number and calls final_answer.
Notice the shape: observe, think, act, observe. If the code raises an error, that is not the end. The error message becomes the next observation, and a capable model usually reads it and fixes its own mistake. Seeing an agent recover from its own bug is a reassuring thing the first time and a very common thing after that.
run(), a variable the code creates in step two is still there in step four. The executor keeps state across steps, so the model can fetch a big table once and slice it several times. This is a real advantage over JSON-style agents, which must pass everything through the conversation.
- Take your task from the previous section and write the loop by hand as a numbered list: what the model would write in step one, what it would observe, what it would write in step two.
- Count how many times the model is called, including the call that produces the final answer.
Installing smolagents and checking the setup
You need Python 3.10 or newer. Always work inside a virtual environment, so that the library and its dependencies stay out of your system Python. On Linux or macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install "smolagents[toolkit]"
On Windows PowerShell the equivalent is below. Use activate.bat in cmd.exe.
py -3.12 -m venv .venv
.venv\Scripts\Activate.ps1
pip install "smolagents[toolkit]"
The quotes around the package name matter. The square brackets select an extra, an optional bundle of dependencies, and shells such as zsh treat unquoted square brackets as a pattern to expand, which produces a confusing "no matches found" error. Quoting avoids that on every shell. If you prefer uv, the same install is uv venv .venv, then activate it, then uv pip install "smolagents[toolkit]".
The core package is deliberately light. The toolkit extra adds the packages behind the built-in web search and page-reading tools. You will meet other extras as you need them, and the library tells you which one when you try to use a feature without it:
openaiforOpenAIModellitellmforLiteLLMModeltransformersfor running models locally (this one pulls in PyTorch, which is large)gradiofor the chat interfacemcpfor connecting to MCP serversdocker,e2b,modalandblaxelfor remote code sandboxes
Verify the install by printing the version:
python -c "import smolagents; print(smolagents.__version__)"
You should see 1.26.0, or a newer number if a release has happened since this was written. The smolagent --help command should also print usage, which confirms the command-line entry point is on your path.
Here is a useful trick for a first session, especially in a classroom with patchy internet or no account yet. You can test your whole setup without any network or API key by giving the agent a fake model that always replies with the same canned answer. This proves the install, the imports and the executor all work, before you add the variable of a real model.
from smolagents import CodeAgent, Model, ChatMessage, MessageRole
class Fake(Model):
def generate(self, messages, stop_sequences=None, **kwargs):
return ChatMessage(
role=MessageRole.ASSISTANT,
content="Thought: done\n<code>\nfinal_answer(6*7)\n</code>",
)
agent = CodeAgent(tools=[], model=Fake(model_id="fake"))
print(agent.run("What is 6 times 7?"))
Running python smoke_test.py prints a banner and some log output, and finally 42. Look closely at what happened: your "model" returned a code block calling final_answer(6*7), the executor ran it, and the loop ended on the first step. You have just watched the entire mechanism work with no intelligence involved.
invalid peer certificate: UnknownIssuer from uv. The fix is to trust your organisation's certificate authority, either with uv pip install --native-tls or by pointing the SSL_CERT_FILE and REQUESTS_CA_BUNDLE environment variables at the CA bundle. Do not turn certificate verification off.
- Create a virtual environment and install
"smolagents[toolkit]". - Print the version and run
smolagent --help. - Run the smoke test above.
- Change the fake model's reply to call
final_answer("hello")and run it again.
42 and later hello. If the smoke test works, any later failure is about the model or the network, not your install.
Choosing a model and giving it credentials
The agent needs a real language model. For a first project the lowest-friction choice is Hugging Face's Inference Providers, which gives you one account and one token that route requests to many hosting companies behind a single interface. A free account includes some credits, enough to follow this guide, and a paid plan raises the allowance.
Create a Hugging Face access token in your account settings and give it the permission to make calls to Inference Providers. Then export it in your shell so the library can find it:
export HF_TOKEN="hf_your_token_here"
On PowerShell the command is $env:HF_TOKEN = "hf_your_token_here". You can also put the line HF_TOKEN=... in a file called .env and load it with python-dotenv, which smolagents already depends on. Never write the token into your source file, and never commit it. A token in a repository is a leaked token.
Now the part that trips almost every beginner, so read it carefully. InferenceClientModel has a default model, and as of this writing that default does not work. The default is a Qwen3 model, but the Hugging Face router currently refuses it with an error saying the requested model is not supported by any provider you have enabled. This is a known open issue in the project. The practical rule is simple: always pass an explicit model_id that you have checked is actually served.
To pick one, look at the router's model list at https://router.huggingface.co/v1/models, or open the page of a model you like on the Hub and check its Inference Providers panel. You want a model that is available through at least one provider and that supports tool calling, because capable instruction-following models write far better agent code than small ones. Weak models are the most common cause of an agent that "does not work". In the examples below, the model name comes from an environment variable so that nothing in the code goes stale when providers change what they host:
import os
from smolagents import InferenceClientModel
model = InferenceClientModel(
model_id=os.environ["MODEL_ID"],
provider="auto",
)
Set MODEL_ID to the identifier you picked, in the same way you set HF_TOKEN. The provider="auto" setting means "use the first provider that serves this model, in the order set in your Hugging Face account settings". You can name a specific provider instead, such as provider="together", if you want predictability.
The same agent works with other model classes. These two lines are the whole difference for OpenAI, and for a local model served by Ollama through LiteLLM:
from smolagents import LiteLLMModel, OpenAIModel
openai_model = OpenAIModel(model_id="your-openai-model-id") # reads OPENAI_API_KEY
local_model = LiteLLMModel(
model_id="ollama_chat/your-local-model",
api_base="http://localhost:11434",
num_ctx=8192,
)
OpenAIModel needs the openai extra and LiteLLMModel needs the litellm extra. The num_ctx=8192 for Ollama matters: Ollama's default context window is very small, and an agent's memory grows quickly, so with the default you will see truncated prompts and bizarre failures. For Gulf and Egyptian employers with data-residency rules, a locally served model or one on a regional cloud endpoint is often the point of choosing these classes, because the prompts and tool outputs then never leave your control. See the Ollama guide if you want to go that route.
InferenceClientModel() with no arguments will fail with model_not_supported while the open issue stands. The command-line tool has the same default and the same problem, so pass --model-id there too.
- Create a Hugging Face token with Inference Providers permission and export it as
HF_TOKEN. - Look at
https://router.huggingface.co/v1/modelsand choose one instruction-tuned model. - Export its identifier as
MODEL_IDand run themodel.pyfile above to confirm it constructs without error.
Your first agent, step by step
Time to build something real. Create a file called first_agent.py. We will use the built-in web search tool, so the agent can look things up, and a CodeAgent, so it writes Python.
import os
from smolagents import CodeAgent, InferenceClientModel, WebSearchTool
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(tools=[WebSearchTool()], model=model)
answer = agent.run(
"How many years passed between the opening of the Suez Canal "
"and the opening of the Aswan High Dam? Show the arithmetic."
)
print(answer)
Run it with python first_agent.py. Here is what you are looking at, in the order it appears, because the console output is the best teacher.
First comes a panel showing the task and the model name. Then a series of steps, each with a header such as Step 1. Inside a step you see the model's reasoning, then a section labelled for the executed code, then the execution logs, which are whatever the code printed, and an output line showing the value of the last expression. Beneath each step there is a footer with the time taken and token counts. The run finishes with a final answer panel.
A representative step looks roughly like this. Your exact words and layout will differ, because the model's wording changes from run to run:
Step 1
Thought: I should look up the opening dates of both projects.
Executing parsed code:
results = web_search("Suez Canal opening year")
print(results)
Execution logs:
## Search Results ...
Notice also that the search tool was called like a normal Python function: web_search("..."). smolagents injects every tool into the executor under the tool's name. The built-in search tool is named web_search. That is why a name collision is possible, and why later you will see errors about duplicate tool names.
If your first run fails, do not panic. The most likely causes are, in order: a missing or wrong HF_TOKEN, a MODEL_ID that is not served, a missing toolkit extra (the error then says you must install ddgs), or a model too weak to produce a valid code block. The error section later in this guide shows the exact messages for each.
- Run
first_agent.pyand read the trace from top to bottom. - Find the step where the agent first called
final_answer. - Count the steps and note the token counts in the footers.
- Run it a second time and compare the two traces.
Writing your own tools
Search is useful, but the real power of an agent is calling code that only you have. In smolagents the simplest way to make a tool is the @tool decorator on an ordinary function. It has strict requirements, and each one exists because the model must be able to read the tool's description.
import os
from smolagents import CodeAgent, InferenceClientModel, tool
@tool
def convert_currency(amount: float, rate: float) -> float:
"""Multiplies an amount of money by an exchange rate.
Args:
amount: The amount in the original currency.
rate: How many units of the target currency one unit of the original buys.
"""
return round(amount * rate, 2)
@tool
def get_office_hours(city: str) -> str:
"""Returns the opening hours of the company office in a city.
Args:
city: The city name, for example Cairo or Riyadh.
"""
hours = {"Cairo": "09:00 to 17:00", "Riyadh": "08:00 to 16:00"}
return hours.get(city, "No office in that city.")
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(tools=[convert_currency, get_office_hours], model=model)
print(agent.run("When does the Riyadh office open? Answer in one sentence."))
Read the three requirements carefully, because breaking any of them gives an error at definition time, before the model is ever called.
Type hints on every argument and on the return value. The library reads amount: float and turns it into a typed input in the description the model sees. Without a hint you get a TypeHintParsingException naming the argument.
A docstring that begins with a plain description. That first sentence is what the model reads to decide whether this tool fits the job. Write it for a smart colleague who has never seen your codebase. Without a docstring you get a DocstringParsingException saying the function has no docstring.
An Args: section documenting every parameter. The parameter descriptions go straight into the model's prompt. If you skip one, the exception says the docstring has no description for that argument. The format is Google style: the argument name, a colon, a description.
The tool's name is the function name, and it must be a valid Python identifier, so convert_currency works and convert-currency does not. The model will write convert_currency(250, 0.5) in its code, so choose names that read well as function calls.
- Write a tool called
word_countthat takes a string and returns the number of words. - Delete the
Args:section and import the file to see the exception. - Restore it, then ask an agent: "How many words are in the phrase 'agents write code'?"
DocstringParsingException naming your argument, then a correct answer of three, and in the trace a call to word_count with the phrase as its argument.
The built-in tools
You do not have to write everything. smolagents ships with a small set of ready-made tools, importable from the top-level package.
WebSearchTool searches the web and returns titles, links and snippets. Its default engine is DuckDuckGo, which needs no account, and you can choose engine="bing" or engine="exa" (the last needs an EXA_API_KEY). The max_results argument limits how many hits come back, and ten is the default. Its tool name is web_search.
VisitWebpageTool fetches a web page and converts it to Markdown text so the model can read it. It truncates what it returns, at 40,000 characters by default, so a huge page cannot flood the model's memory. Its name is visit_webpage.
A common pattern is search followed by reading. Give the agent both tools and it will usually search first, pick a promising link, and open it:
import os
from smolagents import (
CodeAgent,
InferenceClientModel,
VisitWebpageTool,
WebSearchTool,
)
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(
tools=[WebSearchTool(), VisitWebpageTool()],
model=model,
max_steps=8,
)
print(agent.run("Summarise in three sentences what the smolagents library is."))
There is a shortcut you will see in tutorials: add_base_tools=True. It adds the default toolkit for you. For a CodeAgent that means the search and page-reading tools. It is convenient, but be careful. If you pass your own WebSearchTool() and also set add_base_tools=True, the agent ends up with two tools named web_search, and construction fails with ValueError: Each tool or managed_agent should have a unique name! followed by the duplicate names. Pick one approach. Being explicit about the tool list is the better habit anyway: you can see at a glance what your agent is allowed to do.
Search results are untrusted text from the open internet, and the agent will read them. That is a real risk called prompt injection: a web page can contain sentences like "ignore your instructions and do something else", and a model might comply. For now, remember one rule: an agent that reads arbitrary web pages and also runs code deserves more caution than one that does neither. The security section below says what to do about it.
- Run
research_agent.pywith the question changed to something you care about. - Read the trace and write down which URLs it visited.
- Now add
add_base_tools=Truealongside the explicit tools and run it.
ValueError about duplicate names the second. Remove one of the two sources of web_search and it works again.
CodeAgent or ToolCallingAgent
So far every example used CodeAgent. The other class, ToolCallingAgent, has the same constructor but a different way of acting. Understanding the difference helps you choose, and helps you read error messages.
A CodeAgent asks the model to write Python. The tools become functions in that Python. The model can call two tools and combine their results in one block, loop over a list, keep variables and print only what it needs. The price is that you are running model-written code, so the safety question becomes important.
A ToolCallingAgent uses the model provider's built-in tool-calling feature. Each tool is declared to the provider, the model replies with structured requests such as "call get_office_hours with city equal to Riyadh", and smolagents runs exactly that tool with exactly those arguments. The agent cannot do anything you did not wrap in a tool. It can also run several tool calls in the same step in parallel threads, which can save time when the calls are independent. The cost is that every action is one tool call at a time with no variables between them, so tasks that need composition take more steps, and the model and provider must support tool calling properly.
import os
from smolagents import InferenceClientModel, ToolCallingAgent, WebSearchTool
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = ToolCallingAgent(tools=[WebSearchTool()], model=model)
print(agent.run("Who maintains the smolagents library?"))
The official guidance, which this guide follows, is this. Use a CodeAgent when the work involves reasoning, chaining and combining tools, and you can run the code somewhere reasonably safe. Use a ToolCallingAgent for simple, single-purpose tools, when you want strict control over what can happen, and whenever you cannot sandbox code execution. When in doubt as a beginner on a laptop with harmless tools, start with CodeAgent because it is the library's strongest feature, and keep your tools limited.
- Run the same question with a
CodeAgentand aToolCallingAgent, using the same tools. - Compare the number of steps and the total tokens in the traces.
- Ask both a question that needs two tool results combined, such as the difference between two years.
Running code safely: imports and the executor
Because a CodeAgent runs model-written Python, smolagents restricts what that Python may do. Understanding these limits explains a whole family of beginner errors.
The default executor, LocalPythonExecutor, is not Python itself. It is an interpreter written in Python that walks the model's code and evaluates it step by step. It allows only a short list of standard-library modules by default: collections, datetime, itertools, math, queue, random, re, stat, statistics, time and unicodedata. Anything else must be explicitly allowed. If the model writes import os, the run does not execute it. It raises an error and feeds the message back:
InterpreterError: Import of os is not allowed. Authorized imports are: [...]
Many other things are blocked outright, regardless of your settings: eval, exec, compile, access to dunder attributes such as __class__, and modules such as os, subprocess, sys, socket, shutil and pathlib. These blocks are intentional.
To let the agent use a library you trust, name it in additional_authorized_imports. The library must also be installed in your environment.
import os
from smolagents import CodeAgent, InferenceClientModel
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(
tools=[],
model=model,
additional_authorized_imports=["pandas", "numpy.*"],
)
Note the form "numpy.*": it allows the package and all its submodules. A plain "numpy" allows only the top level, and a model that writes import numpy.linalg would be refused. There is also a wildcard "*" that allows everything, and you should never use it casually. Treat every extra import as widening what a model, and anyone who can influence the model, is allowed to do.
If you pass an import that is not installed, the agent fails when you construct it, with InterpreterError: Non-installed authorized modules: .... The same message appears if you mistakenly pass the imports as one string, a trap in the command line tool too. A list of separate strings is correct.
The executor also has guardrails against loops. Code that runs longer than 30 seconds raises an ExecutionTimeoutError, and there are hard caps on the number of operations and while-loop iterations. You can lengthen the timeout through executor_kwargs={"timeout_seconds": 60}.
Now the most important sentence of this section, taken from the project's own security policy: the local executor is not a security boundary. It reduces accidents. It makes it harder for a confused model to delete a directory by mistake. It does not stop a determined attacker, and the project's own history includes published sandbox escapes that were later fixed. Do not run untrusted or prompt-injectable workloads on it. In practice that means: on your own laptop with harmless tools and data you can afford to lose, the local executor is fine for learning. For anything exposed to other people, or any agent that reads untrusted web pages while holding credentials, run the code in a real sandbox, which smolagents supports through Docker, E2B, Modal and Blaxel executors. Setting those up is Mid-level and Senior material, but now you know why they exist.
A version note: older tutorials show a WebAssembly executor (executor_type="wasm"). It was removed in 1.26.0, and asking for it now raises ValueError: Unsupported executor type: wasm.
Import of os is not allowed, resist the urge to allow it. Ask first why the agent wanted it. Usually there is a safer way, such as a tool you write that does the one specific thing needed, with its own checks.
- Ask an agent with no extra imports: "Use the os module to list the current directory."
- Read the error in the trace and see whether the agent finds another way.
- Ask "What is the standard deviation of 2, 4, 4, 4, 5, 5, 7, 9?" and see which module it uses.
os, and a correct answer of two for the second, because statistics and math are allowed by default.
Memory, follow-up questions and passing in data
Each call to agent.run() normally starts fresh. By default it resets memory, so a second run() knows nothing about the first. This surprises beginners who treat the agent like a chat window. If your follow-up refers to "that number", the agent has no idea what number you mean, and its code may fail with The variable x is not defined.
To continue a conversation, pass reset=False:
import os
from smolagents import CodeAgent, InferenceClientModel, WebSearchTool
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(tools=[WebSearchTool()], model=model)
agent.run("What is the capital of Saudi Arabia?")
print(agent.run("Now find its population.", reset=False))
With reset=False the earlier steps stay in memory, so the second run sees them. The trade-off is cost: the memory carries every previous step, so a long conversation sends a large prompt every time.
You can look inside memory after any run. The records are in agent.memory.steps, a list of step objects. The most useful ones are action steps, which hold the model's output, the code that ran, the observations and the token counts. The agent.replay() method pretty-prints the whole run again without calling the model, which is good for reading a long trace calmly.
agent.replay()
for step in agent.memory.steps:
print(type(step).__name__)
print(agent.memory.return_full_code())
The last line concatenates all the code the agent wrote, which is handy when you want to turn a successful agent run into a plain script. To see cost, use agent.monitor.get_total_token_counts(), which returns input, output and total tokens. Older tutorials mention agent.logs and attributes such as total_input_token_count. Those were removed in version 1.21. If a tutorial uses them, it is out of date, even though a page of the official guided tour still mentions agent.logs.
You can also hand data to the agent without pasting it into the prompt. The additional_args argument of run() puts variables straight into the executor, where the model's code can use them by name:
sales = [120, 340, 95, 410, 220]
print(agent.run(
"Using the list called sales, report the total and the average.",
additional_args={"sales": sales},
))
This is better than embedding a long list in the task text. The model still has to be told the variable exists, and that is why the task sentence names it.
If you want structured information about a run, pass return_full_result=True. Then run() returns an object with the output, a state (either success or max_steps_error), the steps, the token usage and the timing, instead of just the answer.
- Run two questions in a row, the second referring to "that", first without and then with
reset=False. - Call
agent.replay()andagent.memory.return_full_code()after the run. - Pass a list through
additional_argsand ask for its maximum.
reset=False and a sensible follow-up with it, the full generated code printed as plain Python, and a correct maximum computed from your list.
The command line and a chat interface
You do not always need a Python file. Installing the package gives you a command called smolagent that runs an agent on one prompt. Remember to pass the model explicitly, for the reason given earlier:
smolagent "What is 17 factorial?" \
--model-type InferenceClientModel \
--model-id "$MODEL_ID" \
--imports math
The flags are mostly self-explanatory. --model-type chooses the class, --model-id the model, --action-type is code (default) or tool_calling, --tools lists tools by name (the built-ins are web_search, visit_webpage and python_interpreter), and --imports lists authorized imports. Both list flags take space-separated values, as in --imports pandas numpy. Do not put them in one quoted string: the official guided tour shows --imports "pandas numpy", which turns into a single module named pandas numpy and fails with Non-installed authorized modules. Running smolagent with no prompt starts an interactive wizard. The CLI reads a .env file in the current directory, and for the key it looks at HF_API_KEY for its --api-key option while the model class itself falls back to HF_TOKEN, so the easiest way to avoid confusion is to rely on HF_TOKEN and not pass --api-key at all.
For a friendlier front end, smolagents can wrap an agent in a Gradio chat page. Install the extra with pip install "smolagents[gradio]" and then:
import os
from smolagents import CodeAgent, GradioUI, InferenceClientModel, WebSearchTool
model = InferenceClientModel(model_id=os.environ["MODEL_ID"])
agent = CodeAgent(tools=[WebSearchTool()], model=model)
GradioUI(agent).launch(share=False)
Pay attention to share=False. The launch method's default is share=True, which creates a public link on gradio.live. Anyone holding that link could talk to an agent that runs code on your machine. Unless you deliberately want to show it to someone remote, and you accept the risk, always pass share=False.
GradioUI.launch() with no arguments shares your agent publicly. For anything running a code executor, treat that as an open door and set share=False.
- Run the
smolagentcommand above with your own question and--imports math. - Run it once with
--imports "math statistics"in quotes and read the error. - Launch the Gradio page with
share=Falseand ask two follow-ups.
InterpreterError about a module named with a space in it, and a local page at an address on your own machine that works for questions but resets memory according to how you set it up.
Configuration you will actually touch
A handful of settings cover most real use. Here is a reference grouped by what you are trying to achieve, all passed to the agent constructor unless stated.
Controlling how long a run can go. max_steps caps the number of action steps; the default is 20. Lower it for simple tasks, so a confused agent stops sooner and costs less. You can also override it for a single call with agent.run(task, max_steps=5).
Controlling how the agent behaves. instructions is a string appended to the built-in system prompt. This is the recommended way to customise behaviour, for example "Answer in formal English and cite the URL you used". Resist rewriting the whole prompt template at first. The library supports it through prompt_templates, but the docs call it generally not advised, and a custom set must contain every required key or you get an AssertionError.
Planning. planning_interval=3 makes the agent write and refresh a plan every three steps, starting with step one. Planning costs extra model calls but can keep long tasks on track. Leave it off until you see an agent wander.
What the model sees. Settings such as temperature and max_tokens go on the model object, and are forwarded on every call. A max_tokens that is too low can cut the model's reply in the middle of a code block, which leads to the code-parsing error. Raising it often fixes mysterious parse failures.
import os
from smolagents import CodeAgent, InferenceClientModel, LogLevel, WebSearchTool
model = InferenceClientModel(
model_id=os.environ["MODEL_ID"],
temperature=0.2,
max_tokens=2048,
)
agent = CodeAgent(
tools=[WebSearchTool()],
model=model,
max_steps=10,
instructions="Be concise. Always mention the source you relied on.",
verbosity_level=LogLevel.INFO,
additional_authorized_imports=["statistics"],
)
- Add
instructionstelling the agent to answer in exactly one sentence, and test it. - Set
max_steps=2and give it a task that needs search plus reading. - Inspect the result of the second run.
Common errors and how to read them
Most beginner trouble falls into a short list. The pattern is the same for each: read the message, decide whether the problem is the model, the setup, or your tool definition, and fix that layer.
Bad request: ... model_not_supported ... not supported by any provider you have enabled. The model you asked for, or the default model you did not override, is not served by the router. It appears wrapped in Error in generating model output. Fix: pass an explicit model_id from the router's list, and check your provider settings.
AgentGenerationError: Error in generating model output: followed by details. This wrapper covers every failure in calling the model: a wrong or missing token (401), used-up credits (402), a timeout, or an unsupported parameter. Always read the nested message after the colon. The wrapper is not the cause.
ModuleNotFoundError: Please install 'litellm' extra to use LiteLLMModel. You used a feature whose extra you did not install. The message includes the exact command, such as pip install 'smolagents[litellm]'. The same pattern exists for openai, transformers, gradio and others.
ImportError: You must install package ddgs to run this tool. You used the search tools without the toolkit extra. Run pip install "smolagents[toolkit]".
AgentParsingError: Error in code parsing: Your code snippet is invalid, because the regex pattern <code>(.*?)</code> was not found. The model did not write a code block in the expected format. The error goes back to the model, which often corrects itself on the next step. If it keeps happening, use a stronger model, raise max_tokens in case replies are truncated, or try code_block_tags="markdown" for models that insist on fenced Markdown blocks.
AgentParsingError: Error while parsing tool call from model output. This comes from a ToolCallingAgent. The model or provider does not support tool calling well. Use a tool-capable model.
InterpreterError: Import of X is not allowed. The model tried an import outside the allowed list. Decide whether to allow it with additional_authorized_imports, or add a safer tool instead.
InterpreterError: The variable x is not defined. Usually the code assumed a variable from an earlier run(), but memory was reset. Use reset=False or additional_args.
ExecutionTimeoutError: Code execution exceeded the maximum execution time of 30 seconds. The code ran too long. Check for an accidental infinite loop, or raise timeout_seconds through executor_kwargs.
Tool definition errors. DocstringParsingException: ... has no docstring, ... no description for the argument 'x', and TypeHintParsingException: Argument x is missing a type hint all mean what they say. Fix the function. Invalid Tool name 'my-tool' means the name is not a valid identifier. Each tool or managed_agent should have a unique name means two tools share a name.
The agent reaches the step limit. The trace ends with Reached max steps, and you still get an answer, but it is a best-effort guess. If you used return_full_result=True, the state is max_steps_error. Fix: look at the trace for where it went around in circles. The cause is usually a vague tool description, a missing tool, or a weak model. Raising max_steps is the last resort, not the first.
Warnings about deprecated names. HfApiModel was renamed InferenceClientModel long ago, and from smolagents import HfApiModel now fails with an ImportError. ManagedAgent is also gone. If a tutorial imports either, it predates the current library.
agent.replay() and ask three questions in order. What did the model actually receive? What did it write? What did the tool or executor return? The fault is almost always at one of those three points.
- Deliberately set
HF_TOKENto a wrong value and run your first agent. - Find the nested message after
Error in generating model output. - Restore the token, then set
max_tokens=20and run again.
Putting it all together
Let us build one small end-to-end project that uses nearly everything above: a price-check assistant for a team that buys cloud services in several countries. It will have two custom tools, one built-in tool, a restricted set of imports, extra instructions, a step limit and a saved record of its code.
The scenario: given a monthly cost in US dollars, the agent must convert it into Saudi riyals and Egyptian pounds using rates we supply, and say which is larger relative to a budget. Because the rates change, the tool takes them from a small dictionary here; in real life that would call your finance system.
import os
from smolagents import CodeAgent, InferenceClientModel, LogLevel, tool
RATES = {"SAR": 3.75, "EGP": 50.0} # illustrative values for the exercise
@tool
def get_rate(currency: str) -> float:
"""Returns how many units of a currency one US dollar buys.
Args:
currency: A three-letter code. Supported codes are SAR and EGP.
"""
if currency not in RATES:
raise ValueError(f"Unsupported currency {currency}. Use SAR or EGP.")
return RATES[currency]
@tool
def check_budget(amount_usd: float, budget_usd: float) -> str:
"""Says whether a monthly cost fits within a monthly budget.
Args:
amount_usd: The monthly cost in US dollars.
budget_usd: The monthly budget in US dollars.
"""
if amount_usd <= budget_usd:
return "within budget"
return f"over budget by {round(amount_usd - budget_usd, 2)} USD"
model = InferenceClientModel(
model_id=os.environ["MODEL_ID"],
temperature=0.2,
max_tokens=2048,
)
agent = CodeAgent(
tools=[get_rate, check_budget],
model=model,
max_steps=8,
additional_authorized_imports=["math"],
instructions=(
"Always fetch rates with get_rate. Never invent a rate. "
"End with a short answer that lists each currency amount."
),
verbosity_level=LogLevel.INFO,
)
result = agent.run(
"Our cloud bill is 1,800 USD a month and the budget is 2,000 USD. "
"Show the cost in SAR and EGP and say whether we are within budget.",
return_full_result=True,
)
print("State:", result.state)
print("Answer:", result.output)
print("Tokens:", result.token_usage.total_tokens)
with open("generated_code.py", "w", encoding="utf-8") as handle:
handle.write(agent.memory.return_full_code())
Walk through what you should see. The agent reads the task, writes code that calls get_rate("SAR") and get_rate("EGP"), multiplies by 1,800, calls check_budget(1800, 2000) and finishes with final_answer containing a sentence. With rates of 3.75 and 50, the amounts are 6,750 riyals and 90,000 pounds, and the status is within budget. The state prints success. The token count tells you what the run cost, which you should compare when you later change the model or the instructions.
The final lines save the code the agent wrote into generated_code.py. Open it. If the agent solved the task the same way every time, you may decide the work does not need an agent at all, and that file is your script. That is a healthy conclusion: agents earn their keep when the path to the answer varies, and a fixed path is better as plain code.
- Run
price_assistant.pyand confirm the numbers by hand. - Ask again with "AED" and read how the agent handles the error.
- Delete the instructions and test whether the agent ever invents a rate.
- Open
generated_code.pyand decide whether this task needed an agent.
What you can now do, and what comes next
You can install smolagents in a virtual environment and verify it, including with a fake model that needs no network. You can connect it to a real model with an explicit model identifier and a token kept in the environment. You can build a CodeAgent or a ToolCallingAgent, give it built-in tools and your own @tool functions, and read the trace to see why it did what it did. You can allow specific imports deliberately, continue a conversation with reset=False, hand data in through additional_args, and pull the generated code out of memory. You know the key errors, and you know which are about the model, the setup or your tool definitions. Most importantly, you know that the local executor is a convenience and not a sandbox, and that the Gradio UI shares publicly unless you say otherwise.
What the Mid-level guide adds: writing tools as classes with setup, loading tools from MCP servers and the Hub, multi-agent systems where a manager delegates to specialists, planning, step callbacks, real sandboxes through Docker and other executors, structured outputs, tracing with OpenTelemetry, and testing agents so you notice when a model change breaks them. The Senior guide covers running agents as a platform: the trust model, concurrency and per-request agents, cost control, upgrades in a library that removes features between minor versions, and when to choose a different tool.
Natural neighbours in this catalogue are MCP, which lets your agent use ready-made tool servers, Langfuse for tracing what agents do, and Hugging Face for the models and Hub behind much of this. Pin your version in your project's requirements, for example smolagents==1.26.0, so that a later release cannot quietly change behaviour under you.
Sources
- smolagents documentation home
- Installation
- Guided tour
- What are agents?
- How do multi-step agents work?
- Building good agents
- Tools tutorial
- Secure code execution
- Inspecting runs
- Agents reference
- Models reference
- Default tools reference
- Python executors reference
- Hugging Face Inference Providers
- smolagents on GitHub
- smolagents releases
- smolagents security policy
- Issue 2584: default model not served