Let an agent delegate self-contained tasks to named child agents.
A single agent that does everything accumulates a large tool set and a long context. Splitting the work across specialized sub-agents keeps each context focused, but wiring up delegation by hand means writing a tool per agent, forwarding deps, threading usage limits, and telling the model what it can delegate to.
SubAgents takes a sequence of SubAgent entries and exposes a single delegate_task(agent_name, task) tool. Each delegation runs the chosen sub-agent in its own run -- with its own message history, so it never sees the parent conversation -- and returns its output to the parent. The available sub-agents are listed in the system prompt as a static instruction, so the listing stays in the cached prefix.
from pydantic_ai import Agent
from pydantic_ai_harness import SubAgent, SubAgents
researcher = Agent('anthropic:claude-sonnet-4-6', name='researcher', description='Researches a topic and reports findings')
writer = Agent('anthropic:claude-sonnet-4-6', name='writer', description='Turns notes into polished prose')
orchestrator = Agent(
'anthropic:claude-opus-4-7',
capabilities=[SubAgents(agents=[SubAgent(researcher), SubAgent(writer)])],
)
result = orchestrator.run_sync('Research the history of TLS and write a one-paragraph summary.')
print(result.output)A delegate's name -- how the parent model refers to it, and how it is listed in the prompt -- is the agent's own name, or a SubAgent(name=...) override. Two delegates resolving to the same name is an error, and an agent with no name and no override is rejected.
| Tool | Purpose |
|---|---|
delegate_task(agent_name, task) |
Run the named sub-agent on a self-contained task and return its output. |
- The sub-agent runs with its own message history, so
taskmust be self-contained. - An unknown
agent_nameraisesModelRetry, so the model can correct itself. - The result returned to the parent is
str(result.output). - With a
modelsmenu configured, the tool takes an extramodelargument (see below).
- Deps are forwarded. The parent run's
depsare passed to each sub-agent, so sub-agents share the parent'sAgentDepsT(enforced by the type signature -- every sub-agent is anAbstractAgent[AgentDepsT, Any]). - Usage is shared by default. The parent's
usageis passed to each sub-agent run, so token usage aggregates and a parentusage_limitsapplies across the whole agent tree. Setforward_usage=Falseto give each sub-agent run its own accounting. - Tools can be inherited. With
inherit_tools=True, the parent agent's own tools (registered directly or viatoolsets) are added to each sub-agent run, on top of the sub-agent's own. Tools contributed by the parent's capabilities are not inherited: they are bound to capability instances registered in the parent run, and would arrive without the hooks and instructions they depend on. Useshared_capabilitiesto give sub-agents a capability. This also excludes the delegate tool itself, so a sub-agent can't recurse into further delegation. Off by default. - Capabilities can be shared.
shared_capabilitiesare applied to every sub-agent run -- e.g. give all sub-agents a common guardrail, memory, or planning capability without rebuilding eachAgent. - Sub-agent events can be streamed. Pass an
event_stream_handlerand it's forwarded to each sub-agent run, so the sub-agent's model-streaming and tool events surface to the caller (the handler receives the sub-agent's ownRunContext).
With include_self=True, the roster also lists the running agent itself, as self. A delegation to self starts a fresh run of that agent (RunContext.agent) on the parent run's model (or the models option the parent picks), with no parent conversation. Because the child is the same Agent, it has every capability, toolset, and instruction bound to it, and they register again in the child run: a guardrail, approval gate, or audit hook bound next to SubAgents sees the tool calls the delegate makes, not only the delegate_task call. This is what Coder uses by default.
from pydantic_ai import Agent
from pydantic_ai_harness import SubAgents
agent = Agent(
'anthropic:claude-opus-4-7',
capabilities=[SubAgents(include_self=True, agent_folders=None)],
)- Only what is bound to the
Agentcarries over. Capabilities, toolsets, instructions, and model settings passed to the parent'srun()are not part of the agent, so the delegate does not get them. PassingSubAgents(include_self=True)itself torun()raises aUserErrorwhen the run starts, since the delegate would come up without it. - Delegation depth is capped. The delegate carries
delegate_tasktoo, somax_depth(default3, counting the top-level run) bounds the tree: the top-level run delegates, its delegates delegate once more, and a run at the limit gets neitherdelegate_tasknor the sub-agent listing. The limit applies to every delegation throughSubAgents, including explicit rosters. - Delegates inherit everything, including what may not suit a sub-task. An
AskUsercapability bound to the agent can prompt the user from inside a delegation, and the delegate returns the agent's ownoutput_type, rendered withstr(). inherit_toolsdoes not apply toself, whose tools are already the parent's. The nameselfis reserved: an explicit delegate with that name is an error, and a disk definition with that name is skipped with a warning.
Each SubAgent carries its own budgets, so one delegate's controls do not touch the others. A SubAgent with no controls set runs with the SubAgents defaults.
from pydantic_ai import Agent
from pydantic_ai.usage import UsageLimits
from pydantic_ai_harness import SubAgent, SubAgents
reproducer = Agent('anthropic:claude-sonnet-4-6', instructions='Reproduce the reported bug from a minimal script.')
librarian = Agent('anthropic:claude-sonnet-4-6', instructions='Find relevant docs, issues, and prior art.')
orchestrator = Agent(
'anthropic:claude-opus-4-7',
capabilities=[
SubAgents(
agents=[
SubAgent(reproducer, usage_limits=UsageLimits(request_limit=35), timeout_seconds=600, max_calls=1),
SubAgent(librarian, usage_limits=UsageLimits(request_limit=18), timeout_seconds=300, max_calls=2),
]
)
],
)| Field | Effect |
|---|---|
models |
Which keys of the SubAgents model menu this delegate may run on, and which one it runs on by default: the first key listed. See "Per-delegation model selection" below. |
usage_limits |
A request/token budget for one delegation. The child runs with its own usage accounting, so the budget counts only that child's requests and tokens (not the parent's or siblings'), even when forward_usage=True. The tradeoff: that child's tokens no longer aggregate into the parent's usage. Reaching the budget is a soft outcome (see below), not a run-stopping UsageLimitExceeded. |
timeout_seconds |
A wall-clock budget for one delegation. When the child exceeds it, its run is cancelled and the parent gets a soft steering message instead of hanging on the child. The cancelled child's event_stream_handler (if any) stops receiving events without a terminal event. |
max_calls |
The maximum number of delegations to this sub-agent per parent run. Once reached, further delegations return a soft budget-exhausted message without running the child. Counts are scoped to one Agent.run (a run_id) and cleared when it ends, so each parent run and each level of a nested tree budgets independently. |
on_failure |
A steering message returned to the parent for any soft degradation of this delegate, in place of the built-in default. Setting it also makes child failures soft (see below). |
contain_errors |
Whether an unexpected crash in this delegate is caught and returned to the parent as a bounded ModelRetry instead of aborting the parent run (see below). Unset inherits the SubAgents(contain_errors=...) default (off). |
The orchestrator knows how hard a task is at the moment it writes the brief, so it is the right place to decide which model runs it. Configure a models menu and the delegate tool gains a model argument that names one of its keys.
from pydantic_ai import Agent
from pydantic_ai.settings import ModelSettings
from pydantic_ai_harness import SubAgent, SubAgents
from pydantic_ai_harness.subagents import ModelOption
reviewer = Agent('anthropic:claude-sonnet-4-6', name='reviewer', description='Reviews a diff')
linter = Agent('anthropic:claude-sonnet-4-6', name='linter', description='Runs the linter and reports failures')
orchestrator = Agent(
'anthropic:claude-opus-4-7',
capabilities=[
SubAgents(
agents=[SubAgent(reviewer), SubAgent(linter, models=['fast'])],
models={
'fast': 'anthropic:claude-haiku-4-5',
'standard': 'anthropic:claude-sonnet-4-6',
'deep': ModelOption(
'anthropic:claude-opus-4-7',
description='hard reasoning, multi-file changes',
settings=ModelSettings(thinking='xhigh'),
),
},
)
],
)- Off by default. With no
modelsmenu themodelargument is not in the tool schema at all, and every delegation runs exactly as it did before. - The keys are the interface. They are listed in the system prompt with each entry's model and description, and the tool schema offers them as an enum, so the model picks from the menu instead of inventing a model name. Name them for the job (
'fast','deep'), not for the vendor. - An entry is a model or a
ModelOption.ModelOption(model, description=..., settings=...)adds a routing hint and per-optionModelSettings, so one key can mean "same model, more thinking". Those settings merge over the sub-agent's ownmodel_settings, which keeps the parts the option does not set. - Resolution order. The key the parent passed, else the delegate's first allowed key (see below), else the delegate's own model, else the parent run's model.
- A delegate can be restricted.
SubAgent(linter, models=['fast'])pins that delegate tofast: the first listed key is what it runs on when the parent passes nomodel, and the others are refused with aModelRetry. The restriction is rendered in the prompt listing (- linter: Runs the linter (models: fast)). Restricting to a key the menu does not define is aValueErrorat construction. - A rejected key costs nothing. An unknown or unavailable key comes back as a
ModelRetrylisting the valid options, before the delegate'smax_callsbudget is charged.
A soft outcome returns a steering message to the parent as a normal tool result, so its model reads the message and decides what to do next (rather than immediately re-delegating, which a ModelRetry invites). A timeout, a reached usage_limits budget, and an exhausted max_calls budget are always soft. When on_failure is set, the message it carries replaces the built-in default for these outcomes.
A sub-agent run that fails with a soft model error (ModelRetry, UnexpectedModelBehavior, e.g. it exhausted its own retries) is, by default, converted into a ModelRetry for the parent -- so the parent's model sees Sub-agent '<name>' failed: ... and can react by re-delegating. The delegate tool defaults to tool_retries=2, so the parent aborts only after that many consecutive delegate failures; the counter resets after any successful delegation. Raise tool_retries to tolerate a flakier sub-agent, or set None to inherit the parent agent's default tool retries. Set on_failure for a delegate to make its failures soft instead: the child error returns the on_failure message as a normal tool result.
Hard errors propagate to stop the whole run. A UsageLimitExceeded from a child that has no per-delegate usage_limits (so it shares the parent's accounting) means the whole tree is out of budget and propagates; a child reaching its own usage_limits is soft, as above.
An unexpected crash -- any other exception the child raises, such as a provider ModelAPIError/FallbackExceptionGroup or a plain ValueError from a bad tool argument -- propagates by default and aborts the parent run. Set contain_errors=True (per delegate, or as the SubAgents default) to catch it and return it to the parent as a bounded ModelRetry instead, so one delegate crash cannot kill the whole run. Containment stays loud: the exception rides the retry message (Sub-agent '<name>' crashed: ...), it is logged via the standard logging module, and tool_retries still bounds consecutive crashes into an abort. This is orthogonal to on_failure -- a contained crash always raises the loud retry, never the soft on_failure return, so a genuine bug is never masked as success. Cancellation, a shared UsageLimitExceeded, pydantic-ai control-flow signals (CallDeferred, ApprovalRequired, the Skip* signals), and UserError bypass containment regardless of contain_errors. Cancellation covers both kinds: external cancellation (asyncio.CancelledError) propagates as-is, and a child's own first-party cancellation (RunContext.cancel() inside the child, raising RunCancelled) leaves the delegate tool uncontained, after which pydantic-ai isolates it as a failed delegate_task return the parent model can react to, not a crash retry that invites re-delegation.
SubAgents emits typed capability events in the sub_agents namespace so a host can show a delegation as it runs, and how it ended, without parsing the delegate tool's arguments and result:
| Event | Dispatch | When | Payload |
|---|---|---|---|
DelegationStartEvent |
stream | a delegation passed every check and the child run is about to start | agent_name, task, truncated, model (the menu key, or None), inherits_tools |
DelegationEndEvent |
stream | the delegation settled into what the parent receives | agent_name, outcome, output, truncated, usage, duration_seconds |
Both are notifications. One delegation is one delegate_task call, so a start and its end share the tool_call_id core stamps on every event; that is how a subscriber pairs them when the model delegates in parallel.
outcome follows the failure handling above: ok (the child's output went back to the parent), timeout, budget (the child's own usage_limits), failed (a soft model error, returned as on_failure or raised as a ModelRetry), or contained (a crash contain_errors caught). output is what the delegate tool hands back to the parent in each case: the child's output, the steering message, or the retry text. It is emitted from inside the tool, before any ToolGuardrail result guard screens that text for the model; a host that needs the screened version reads the ToolReturnPart in core's FunctionToolResultEvent. usage is the child's own RunUsage when it has separate accounting (usage_limits set, or forward_usage=False) and None when it accrues into the parent's usage, where its share is not separable.
A delegation refused before the child runs (an unknown sub-agent, a model key off the menu, an exhausted max_calls budget) emits nothing; the tool result says why. An exception that propagates out of the delegate tool (a shared usage limit, an uncontained crash, a cancellation) ends without an end event. A SubAgentToolset registered directly in Agent(toolsets=[...]) has no owning capability and emits nothing. As with any capability event, a listener that raises aborts the parent run.
task and output are cut at MAX_EVENT_TEXT_CHARS (4096) with a truncated flag, so a persisted or forwarded event stream cannot be flooded by one verbose delegation.
from pydantic_ai import Agent
from pydantic_ai.capabilities import AbstractCapability, on_event
from pydantic_ai_harness.subagents import DelegationEndEvent, SubAgent, SubAgents
class ReportDelegations(AbstractCapability):
@on_event(DelegationEndEvent)
async def on_delegation_end(self, ctx, event: DelegationEndEvent) -> None:
print(f'{event.agent_name}: {event.outcome} in {event.duration_seconds:.1f}s')
researcher = Agent('anthropic:claude-sonnet-4-6', name='researcher')
agent = Agent(
'anthropic:claude-opus-4-7',
capabilities=[SubAgents(agents=[SubAgent(researcher)]), ReportDelegations()],
)Nested model streaming from the child run is not an event concern; pass an event_stream_handler for that.
SubAgents emits no OpenTelemetry spans of its own: the child run is a core agent run with its own spans nested under the parent's tool-call span, and the events above carry the outcome a trace would only show as an exception or a tool result.
The sub-agents are listed in the system prompt via get_instructions, using each agent's description (or a SubAgent(description=...) override). A sub-agent with no description is listed by name alone.
A repo's markdown agent definitions become delegates without writing any Agent code. By default every *.md file under the conventional folders is loaded as a sub-agent, alongside the explicitly-passed agents.
from pydantic_ai import Agent
from pydantic_ai_harness import SubAgents
orchestrator = Agent(
'anthropic:claude-opus-4-7',
capabilities=[SubAgents(inherit_tools=True)], # auto-loads ./.agents/agents/ and ~/.agents/agents/
)agent_folders controls where definitions come from. It defaults to 'agents', the conventional layout:
- A folder-name
str(the default'agents'): for the project root (cwd) then the home root, load from<root>/.agents/<name>/, falling back to<root>/.claude/<name>/when<root>/.agents/is absent. - A sequence of paths loads from exactly those folders, in order.
Nonedisables disk loading, exposing only the explicitly-passedagents.
A definition is a markdown file with optional frontmatter:
---
name: researcher
description: Researches a topic and reports findings
tools: Read, Grep
---
You research topics. Report your findings, each with a source.nameis the delegate name (how the parent refers to it and how it is listed). It falls back to the filename stem when absent.descriptiondrives the prompt listing.- The markdown body becomes the agent's instructions.
tools(orallowed-tools) is a comma-separated string or a YAML block list. See "Tools" below.modelandcolorare ignored: the model is inherited from the parent (see below), andcolorhas no pyai equivalent.
Frontmatter is read by a small, dependency-free parser limited to those keys (pyyaml is not a harness dependency). Full YAML frontmatter is not supported.
Disk agents inherit the parent run's model by default. Per agent, the caller can override the model and set a thinking/effort level via agent_overrides, keyed by the agent's name:
from pydantic_ai_harness import SubAgents
from pydantic_ai_harness.subagents import AgentOverride
SubAgents(
agent_folders='agents',
agent_overrides={'researcher': AgentOverride(model='anthropic:claude-sonnet-4-6', effort='high')},
)Every agent the capability builds runs at a minimum thinking-effort floor. MINIMUM_EFFORT_FLOOR and the clamp_effort(level, floor=...) helper are exported so an orchestrator can apply the same floor to its own agents (that orchestrator-side application is the caller's responsibility). clamp_effort maps None/False to the floor, leaves True (provider-default effort) unchanged, and raises a concrete level below the floor up to it. Effort is applied through pyai's ModelSettings.thinking.
A disk agent gets no tools by default (inherit_tools is False); set inherit_tools=True to expose the parent's tools to it through the inherit_tools mechanism, in which case its tools frontmatter is ignored. To map the frontmatter tool names to specific toolsets instead, pass a tool_resolver: it receives each tool name (so it can honor entries like Bash(git:*)) and returns the toolsets that provide it, or None for an unknown name, which is skipped with a warning.
from pydantic_ai_harness import SubAgents
def resolve(tool_name: str):
return TOOLSETS.get(tool_name) # -> Sequence[AgentToolset[object]] | None
SubAgents(agent_folders='agents', tool_resolver=resolve)When the same name appears in more than one source, the higher-precedence one wins and the others are skipped with a warning: explicitly-passed agents first, then the project folder, then the home folder (and, for an explicit path sequence, earlier paths before later ones). A duplicate name within the explicitly-passed agents list is still an error.
SubAgents(
agents=(), # Sequence[SubAgent[AgentDepsT]] -- each pairs an agent with its run controls
models={}, # Mapping[str, Model | str | ModelOption] -- per-delegation model menu (off when empty)
agent_folders='agents',# folder-name str (convention) | Sequence[Path] | None (disable)
agent_overrides={}, # Mapping[str, AgentOverride] -- per-disk-agent model/effort override
tool_resolver=None, # Callable[[str], Sequence[AgentToolset[object]] | None] -- disk-agent tool mapping
forward_usage=True, # share the parent's usage with sub-agent runs
inherit_tools=False, # expose the parent's own tools to sub-agents (capability tools excluded)
shared_capabilities=(),# capabilities applied to every sub-agent run
event_stream_handler=None, # forwarded to each sub-agent run to stream its events
tool_name='delegate_task',
tool_retries=2, # extra delegate-tool attempts after a sub-agent error before aborting (None inherits the agent default)
contain_errors=False, # default for SubAgent.contain_errors: contain an unexpected crash as a bounded retry
include_self=False, # also list the running agent itself as the delegate `self`
max_depth=3, # delegation levels, counting the top-level run
)SubAgent(
agent, # AbstractAgent[AgentDepsT, Any] -- the child agent to run
name=None, # delegate name; defaults to the agent's own `name`
description=None, # prompt-listing description; defaults to the agent's own `description`
models=None, # Sequence[str] -- menu keys this delegate may run on; the first is its default
usage_limits=None, # per-delegation request/token budget (isolated accounting)
timeout_seconds=None, # per-delegation wall-clock budget
max_calls=None, # max delegations to this sub-agent per parent run
on_failure=None, # steering message for soft degradations of this delegate
contain_errors=None, # contain an unexpected crash as a bounded retry; None inherits the SubAgents default
)SubAgents is not serializable via the agent spec (it holds live Agent instances), so get_serialization_name() returns None.
- Sub-agents can themselves have
SubAgents, forming a tree. EachSubAgentsstops offering delegation once the run is at its ownmax_depth, so a delegate with a higher limit can go deeper than its parent's. Shareusage(the default) and set ausage_limitson the top-level run to bound the whole tree. - Delegations the model issues in parallel run as independent sub-agent runs.