smolagents
TL;DR
smolagents is a compact Python agent library from Hugging Face whose main CodeAgent expresses each action as a Python snippet instead of a JSON tool call. It supports manager agents that delegate to named sub-agents, MCP tools and remote sandboxes. It suits Python developers who want a small, readable codebase over a large framework.
Key facts
| Type | Framework |
|---|---|
| Languages / SDKs | Python |
| License | Apache-2.0 |
| Pricing model | Open source, free |
| Orchestration pattern | Supervisor |
| GitHub stars | 29,614 (as of 2026-10-01) |
| GitHub forks | 3,017 |
| Last push | 2026-09-30 |
| Latest release | v1.26.0 |
| Repository | huggingface/smolagents |
| Website | huggingface.co |
| Documentation | huggingface.co |
| Last verified | 2026-09-30 |
Key features
- Two agent types: CodeAgent writes each step as Python code, ToolCallingAgent emits structured JSON tool calls. (source)
- Multi-agent setups by passing named agents with a description to a manager agent's managed_agents argument. (source)
- MCPClient loads tools from stdio or Streamable HTTP MCP servers, including structured tool output. (source)
- Agent-generated code can run in Blaxel, E2B, Modal or Docker sandboxes via executor_type. (source)
- Agents and tools can be pushed to and loaded from the Hugging Face Hub; GradioUI gives a chat front end. (source)
- Runs are traced through OpenTelemetry instrumentation, with Phoenix and Langfuse examples. (source)
- Step memory can be replayed with agent.replay() and edited in place between steps. (source)
Architecture and orchestration pattern
Pattern: Supervisor
Each agent runs a ReAct-style loop: the model reads the task plus prior steps, proposes an action, the framework executes it and appends the observation, and the loop stops when the agent calls final_answer or hits max_steps. In a CodeAgent the action is a Python snippet run by an interpreter or remote executor; in a ToolCallingAgent it is a JSON tool call. An optional planning_interval inserts periodic planning steps.
Multi-agent work is hierarchical. A sub-agent gets a name and description and is passed to a manager through managed_agents; the manager then calls it much like a tool and receives its result. There is no shared blackboard or graph definition.
State lives in agent.memory, an ordered list of typed steps (system prompt, task, planning, action). Calling run(..., reset=False) keeps that memory for a follow-up run. No built-in long-term or cross-session store is documented.
Human in the loop
The docs show human review through step_callbacks: a callback registered for PlanningStep pauses the run after a plan is written so a person can approve, edit or cancel it, and run(task, reset=False) resumes with memory intact. A running agent can be stopped with agent.interrupt() (the guided tour suggests wiring it to a button in GradioUI), and memory steps can be edited before the next step. No built-in per-tool approval gate is documented.
Protocols
| Protocol | Support | Note |
|---|---|---|
| MCP | Yes evidence | Client support: MCPClient (install the mcp extra) loads tools from stdio and Streamable HTTP MCP servers into an agent. |
| A2A | Unknown | Searched README, docs source (docs/source/en), GitHub code search and issues for a2a / agent2agent; nothing official found. |
| AG-UI | Unknown | Searched README, docs source, GitHub code search for ag-ui / agui, and the AG-UI README integration list; smolagents is not listed. |
Best for
- Python developers who want agents that compose several tool calls inside one code action.
- Web research agents built as a manager plus search sub-agent, as in the official multi-agent example. (shortlist)
- Running agents on local models through TransformersModel or Ollama via LiteLLM. (shortlist)
- Learning how an agent loop works from a small codebase (agents.py is under 1,000 lines per the README).
Not for
- Teams that need a TypeScript, Java or .NET SDK.
- Executing untrusted model-written code without setting up a remote or container sandbox.
- Designs that need explicit graphs, persisted checkpoints or long-term memory out of the box.
Quickstart
pip install "smolagents[toolkit]" from smolagents import CodeAgent, ToolCallingAgent, InferenceClientModel, WebSearchTool
model = InferenceClientModel() # uses HF_TOKEN from the environment
searcher = ToolCallingAgent(
tools=[WebSearchTool()],
model=model,
name="searcher",
description="Looks things up on the web and reports short findings.",
max_steps=6,
)
lead = CodeAgent(tools=[], model=model, managed_agents=[searcher])
print(lead.run("In which year was Python 3.0 released, and how many years before 2026 was that?"))
Common pitfalls
- Requires Python 3.10 or newer.
InferenceClientModelneeds anHF_TOKENfor Hugging Face Inference Providers; other providers need their own keys.- Managed agents must have both
nameanddescription, or the manager cannot call them. CodeAgentonly allows whitelisted imports; add more withadditional_authorized_imports.- The default
LocalPythonExecutoris not a security sandbox; use Blaxel, E2B, Modal or Docker for untrusted code. - MCP tools need the extra:
pip install "smolagents[mcp]". - With Ollama through
LiteLLMModel, the guided tour warns that Ollama's default 2048-token context fails and setsnum_ctx=8192.
Pros
- Small core: the README states the main agent logic in agents.py is under 1,000 lines, which keeps it readable and hackable. (source)
- Model-agnostic: model classes cover HF Inference Providers, LiteLLM, OpenAI-compatible servers, local transformers, Azure and Bedrock. (source)
- Four documented remote or container sandboxes (Blaxel, E2B, Modal, Docker) for running generated code. (source)
- Documented human-in-the-loop pattern for approving or editing plans mid-run. (source)
- OpenTelemetry-based tracing with worked examples for Phoenix and Langfuse. (source)
Cons
- The built-in LocalPythonExecutor is explicitly not a security boundary; an open issue reports a sandbox escape via ctypes. (source)
- Sandboxing only the code snippets (executor_type) does not support multi-agent setups, per the secure execution guide. (source)
- An open bug reports managed-agent hierarchies with two levels not working. (source)
- An open bug reports managed agents sharing state when run in parallel. (source)
- Python only; no official SDK in other languages is documented. (source)
Alternatives
FAQ
Does smolagents support MCP?
Yes, as a client. With the mcp extra installed, MCPClient loads tools from stdio or Streamable HTTP MCP servers and hands them to an agent. A2A and AG-UI support were not found in the official docs.
Is smolagents free?
The library is Apache-2.0 licensed with no paid tier of its own. Model calls are billed by whichever provider you configure; Hugging Face Inference Providers need an HF_TOKEN.
What language is smolagents written in?
Python. The installation guide requires Python 3.10 or newer, and no SDK for other languages is documented.
How do multiple agents work together in smolagents?
Through a hierarchy: sub-agents with a name and description are passed to a manager agent via managed_agents, and the manager calls them as it would call a tool.
Is it safe to let a CodeAgent run code locally?
The project warns that LocalPythonExecutor is not a security boundary. For untrusted code the docs recommend Blaxel, E2B, Modal or Docker executors.
Sources
- smolagents GitHub repository (README)
- smolagents documentation
- Installation options
- Guided tour
- Tools tutorial (MCPClient)
- Orchestrate a multi-agent system
- Human-in-the-loop plan customization
- Manage your agent's memory
- Secure code execution
- Inspecting runs with OpenTelemetry
- Issue #1061: two-level managed agent hierarchy
- Issue #1781: managed agents share state in parallel runs
- Issue #2094: sandbox escape via ctypes in LocalPythonExecutor