Best multi-agent frameworks for research agents
Research agents search, read, cross-check and write reports, usually by splitting a question across workers and merging the results. This page lists frameworks and harnesses whose records document that pattern. The order follows the criteria below. Data as of 2026-09-30.
Selection criteria
- Documents a way to split a task across agents and merge results (supervisor, handoff, workforce or graph).
- Can persist or resume long runs.
- Connects to external tools and data (MCP evidence on the tool page).
- Runs on your own infrastructure with an open-source license.
- Supports human review of plans or tool calls.
Shortlist
- DeerFlow: MIT-licensed harness built on LangGraph for research, reports and coding: a lead agent spawns sub-agents with their own context,
batch_taskruns large item sets from a durable SQL-backed queue, and MCP servers are configurable. Needs about 4 vCPU and 8 GB RAM per its README. - LangGraph: Graph runtime with checkpoints at every step,
interrupt()for human review and time travel, which suits long research runs that must resume; multi-agent patterns are composed from subgraphs. It is low-level by its own description. - Deep Agents: LangChain's agent harness on LangGraph with subagents in isolated context, a pluggable filesystem and tool-call approval, so a long-horizon research agent does not have to be assembled from parts.
- LlamaIndex:
AgentWorkflowhandoffs on event-driven Workflows plus a large catalog of data integrations (the README counts over 300 packages) and a serializable runContext. The README notes the company's focus has moved to LlamaParse. - CAMEL: Apache-2.0 research framework whose Workforce splits tasks across worker agents; the official example is a searcher, analyst and writer team. A2A and AG-UI are unknown.
- Haystack: Explicit retrieval pipelines with an
Agentcomponent and a coordinator-plus-specialists pattern, breakpoints with JSON snapshots, and tool-level human confirmation. - CrewAI: Role-based crews run sequentially or under a manager, with checkpointing and
@human_feedbackin Flows. Python >=3.10,<3.14.
Shortlist facts
| Tool | Type | Languages | License | Pricing | Pattern | MCP | A2A | AG-UI | Stars |
|---|---|---|---|---|---|---|---|---|---|
| DeerFlow | Harness | Python, TypeScript | MIT | Open source, free | Supervisor | Yes evidence | Unknown | Unknown | 83,261 |
| LangGraph | Framework | Python | MIT | Open core | Graph | Yes evidence | Yes evidence | Partial evidence | 42,511 |
| Deep Agents | Framework | Python, TypeScript | MIT | Open core | Supervisor | Yes evidence | Partial evidence | Partial evidence | 29,870 |
| LlamaIndex | Framework | Python | MIT | Open core | Handoff | Yes evidence | Unknown | Yes evidence | 52,370 |
| CAMEL | Framework | Python | Apache-2.0 | Open source, free | Supervisor | Yes evidence | Unknown | Unknown | 17,800 |
| Haystack | Framework | Python | Apache-2.0 | Open core | Supervisor | Yes evidence | Unknown | Unknown | 26,633 |
| CrewAI | Framework | Python | MIT | Open core | Crew / roles | Yes evidence | Yes evidence | Partial evidence | 59,217 |
FAQ
What is the difference between a research harness and a framework?
DeerFlow and Deep Agents ship a ready agent with sub-agents, sandbox or filesystem and memory. LangGraph, LlamaIndex, CAMEL, Haystack and CrewAI are libraries from which you build the research flow yourself.
Which of these can run fully on my own machine?
All seven are open-source libraries or self-hosted apps. Model access is separate: each record lists whether local models (for example through Ollama) are documented.
Do they support MCP for search and data tools?
Yes. Every tool on this list has an MCP cell marked yes with an evidence link on its page.