Skip to content
CCAR-FAcademy
Domain 5 · Statement 5.3 3 of 6
5.3

Implement error propagation strategies across multi-agent systems

  • A propagated error must carry failure type, what was attempted, partial results and suggested alternatives — that is what makes coordinator recovery intelligent rather than blind.
  • Distinguish access failures (source never consulted; a retry decision applies) from valid empty results (query succeeded, nothing matched; that is evidence).
  • Generic statuses like "search unavailable" are a distractor: they discard exactly the context the coordinator needs.
  • Silent suppression and whole-workflow termination are both wrong; subagents recover transient failures locally and propagate only the unresolved ones.
  • Synthesis output needs coverage annotations so gaps from unavailable sources are visible instead of reading as findings.

In a multi-agent system a subagent failure is information, and the reliability question is how much of that information survives the trip back to the coordinator.

Three outcomes, three recovery branches. The taxonomy is the statement. Label which outcome occurred and the coordinator knows what to do; collapse them into one status and it cannot.

Subagent outcome What it means Who resolves it What the coordinator does
Transient access failure — timeout, rate limit, 500, auth rejection The source was never consulted, and the fault may clear on its own The subagent, locally: bounded retry, backoff, a narrower query Nothing; it never sees the ones that clear
Unresolved access failure Local recovery is exhausted and the topic is still uncovered Propagated upward, with its context attached Retry, re-route to an alternative source, narrow the scope, or record a gap
Valid empty result The query executed successfully and nothing matched Nobody — there is nothing to repair Accept it as evidence and treat the question as answered

An empty result is not a broken one. This is the distinction the exam leans on hardest, and it is easy to lose because both outcomes look alike in a log line. Absence is itself a finding: it is what the system now knows. Retrying it spends budget on a settled question, and re-reporting it as a failure understates the coverage the run actually achieved. The reverse error is worse. An access failure dressed as an empty result puts a silent hole into the evidence base, and nothing downstream can tell it apart from a real answer. The MCP specification draws the same boundary at the protocol level, separating a protocol error from a tool execution error reported with isError: true.22

What a propagated failure has to carry. Four things: the failure type, what was attempted (the actual query or target), any partial results already obtained, and alternative approaches the subagent can suggest.23 Four designs compete for the job, and three of them destroy part of that payload:

Design What reaches the coordinator What it costs
Structured report upward Failure type, attempted query, partial results, suggested alternatives Nothing — it is the only design that supports an intelligent recovery decision
Generic status such as "search unavailable" One opaque string The coordinator cannot tell whether retrying is sensible, whether half the work is already done, or whether another source would answer
Empty results returned as success Nothing; the gap is invisible Synthesis draws confident conclusions over data that was never retrieved
Terminate the whole workflow The failure, and nothing else Every other subagent's completed work is thrown away over one recoverable problem

The correct posture sits between silence and shutdown: resolve transient failures where they happen, and propagate only what you cannot resolve yourself.26

Failures must reach the final artifact. Recovery is not complete when the coordinator handles the error; it is complete when the report is honest about it. Structure synthesis output with coverage annotations that distinguish well-supported findings from topic areas with gaps caused by unavailable sources. A reader who cannot see which sections rest on unavailable evidence will treat a hole as a conclusion.

Error propagation from subagent to synthesisTrace one failure end to end: local recovery first, then a structured propagation that lets the coordinator re-route, and finally a coverage annotation in the output.CoordinatorCoordinatorWeb search subagentWebsearchsubagentDocument subagentDocumentsubagentSynthesis agentSynthesisagentresearchsubtopic Alocal retry,transient timeoutstructured error pluspartial resultstry suggestedalternative sourcevalid empty result,query succeededfindings plusknown gapsreport with coverageannotations
Error propagation from subagent to synthesis

Trace one failure end to end: local recovery first, then a structured propagation that lets the coordinator re-route, and finally a coverage annotation in the output.

Access failure vs. valid empty resultCement the distinction that drives most items here, and show what each of the two anti-patterns costs the coordinator.Access failureAccess failureValid empty resultValid emptyresultRetry or reroute decisionRetry or reroutedecisionAccept as evidenceAccept asevidenceAccess failure disguised as emptyAccess failuredisguised asemptyFalse confidence in reportFalse confidencein reportEmpty reported as generic errorEmpty reported asgeneric errorWasted retries, lost contextWasted retries,lost contextcorrect, source neverconsultedanti-pattern, silentlysuppressedgap reads as findingcorrect, query ran and matchednothinganti-pattern, evidencediscardedcoordinator flying blind
Access failure vs. valid empty result

Cement the distinction that drives most items here, and show what each of the two anti-patterns costs the coordinator.

A structured subagent error contract

Scenario 3 · Multi-Agent Research System

The web-search subagent hits repeated timeouts on one provider. It retries locally with backoff, and only after exhausting its local budget does it return a structured failure that keeps everything the coordinator needs — including the two results it did manage to fetch and an alternative it cannot itself reach.

typescript
type SubagentResult<T> =
  | { status: 'ok'; data: T[] }
  | { status: 'empty'; query: string; note: string }   // query ran, no matches
  | {
      status: 'failed';
      failureType: 'timeout' | 'rate_limited' | 'auth' | 'not_found' | 'malformed';
      attempted: string;            // the exact query or target
      localRecovery: string;        // retries already performed
      partialResults: T[];          // keep what was obtained
      alternatives: string[];       // routes the coordinator could take
    };

const result: SubagentResult<Finding> = {
  status: 'failed',
  failureType: 'timeout',
  attempted: 'web_search("EU battery directive 2026 compliance costs")',
  localRecovery: '3 attempts, exponential backoff, then narrowed query',
  partialResults: [finding1, finding2],
  alternatives: ['regulator PDF archive via document_agent', 'query without year filter'],
};
Discriminated result type: success, empty, or failure — never conflated.

The taxonomy earns its keep at the coordinator

Scenario 6 · Structured Data Extraction

An extraction worker looks up a supplier VAT number in the reference service. "Not registered" and "reference service unreachable" must not reach the coordinator as the same status: the first is a finding the validator can act on, the second leaves the field unverified and needs a retry decision. Collapsing them means unverified records are silently promoted as verified-absent. Note the division of labor — the shape of the error payload a tool hands back is a tool-design question (2.2); what matters here is that the coordinator's recovery branch is genuinely different for the two outcomes, and that the difference survives into the record.

text
subagent outcome            coordinator recovery branch
-----------------------------------------------------------------------
valid empty result          accept as evidence; record "not registered";
(query ran, zero matches)   do NOT retry; the field is complete

access failure              field stays unverified; choose retry, alternate
(timeout, 429, 5xx, auth)   source, or human review; annotate the record as
                            uncovered — never mark the field absent

both flattened into         coordinator cannot tell them apart; it either
"lookup unavailable"        wastes retries on a settled question or promotes
                            unverified records as verified-absent
One taxonomy, two recovery branches — and what collapsing them costs.

Coverage annotations in the synthesized report

Scenario 3 · Multi-Agent Research System

Two of five sources were unreachable. The report neither omits the affected topic nor states its conclusions with the same force as the rest. An explicit coverage section tells the reader which claims are well-supported and which topic areas have gaps due to unavailable sources — the difference between a limitation and a silent error.

markdown
## Coverage and confidence

**Well-supported** (3+ independent sources, all reachable)
- Adoption growth in the EU market, 2024-2026.
- Regulatory timeline for the 2027 compliance deadline.

**Partially covered** (single source, or partial results only)
- Cost-per-unit impact: 2 of 6 planned sources retrieved before timeout.

**Gaps due to unavailable sources**
- FY2025 filings: document store returned auth failures on 2 attempts.
  No conclusion is drawn about FY2025 in this report.
- Competitor pricing: search provider rate-limited; retry scheduled.
Report section that makes gaps legible.
  • Returning a generic status such as "search unavailable" instead of failure type, attempted query, partial results and alternatives because the coordinator then has nothing to base a retry-or-reroute decision on.
  • Reporting an empty result as a failure (or a failure as an empty result) instead of distinguishing them because one wastes retries on settled questions and the other lets unverified gaps pass as evidence.
  • Silently suppressing a subagent error and returning empty results as success because the coordinator will synthesize confident conclusions over data that was never retrieved.
  • Terminating the whole workflow on a single subagent failure instead of recovering locally and propagating only unresolved errors because it discards every other subagent completed result for a recoverable problem.
  • Look for stems where a subagent returns "no results" and the coordinator concludes the topic is settled. The credited answer separates access failures from valid empty results.
  • Options that "fail fast and abort the workflow" and options that "return an empty array so the pipeline continues" are usually the paired distractors — the guide names both as anti-patterns.
  • When the stem is about the final deliverable rather than the run, the answer is usually coverage annotations distinguishing well-supported findings from gaps, not a retry policy.
References — 3 sources
  1. Tools Model Context Protocol The normative source for `isError`. It splits failures into protocol errors (JSON-RPC) and tool execution errors (`isError: true`), and states the reason this unit's whole argument depends on: "Clients SHOULD provide tool execution errors to language models to enable self-correction." It also fixes the result shape — model-visible data travels in `content` or `structuredContent`, and a tool returning structured content "SHOULD also return the serialized JSON in a TextContent block".
  2. Connect to external tools with MCP Anthropic The SDK's error-handling section, including per-server status reporting (`mcpServerStatus()` / `get_mcp_status()`) — the difference between a tool that failed and a server that never connected, which the four error categories do not cover.
  3. Client Best Practices Model Context Protocol The spec's own quantified threshold for "too many tools": switch to progressive discovery once tool definitions occupy 1–5% of the context window. It turns the unit's unthresholded scoping rule into a number.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • Structured error context (failure type, attempted query, partial results, alternative approaches) as enabling intelligent coordinator recovery decisions
  • The distinction between access failures (timeouts needing retry decisions) and valid empty results (successful queries with no matches)
  • Why generic error statuses ("search unavailable") hide valuable context from the coordinator
  • Why silently suppressing errors (returning empty results as success) or terminating entire workflows on single failures are both anti-patterns

Skills in

  • Returning structured error context including failure type, what was attempted, partial results, and potential alternatives to enable coordinator recovery
  • Distinguishing access failures from valid empty results in error reporting so the coordinator can make appropriate decisions
  • Having subagents implement local recovery for transient failures and only propagate errors they cannot resolve, including what was attempted and partial results
  • Structuring synthesis output with coverage annotations indicating which findings are well- supported versus which topic areas have gaps due to unavailable sources
Back to top