Skip to content
CCAR-FAcademy
Domain 2 · Statement 2.2 2 of 5
2.2

Implement structured error responses for MCP tools

  • MCP signals failure with the isError flag, but the flag alone carries no recovery information — the structured metadata you attach does.
  • Return errorCategory (transient / validation / permission), an isRetryable boolean, and a human-readable description on every failure.
  • Business rule violations get retriable: false plus a customer-friendly explanation the agent can relay, so it stops retrying and communicates the policy. The guide spells the flag both isRetryable and retriable — recognize either.
  • Uniform "Operation failed" responses force the agent to guess, which produces wasted retries or premature give-up.
  • Subagents recover from transient failures locally and propagate to the coordinator only what they cannot resolve — with the four elements of structured error context: failure type, what was attempted, partial results, and alternative approaches.
  • A successful query with no matches is a valid empty result, not an error; conflating it with an access failure triggers pointless retries.

An error is a message to a reasoning system

MCP communicates tool failure back to the agent with the isError flag on the tool result. That flag says only "this call failed" — it carries no recovery information by itself. Everything the agent needs in order to choose a next action has to be in the payload you attach to it.22

This is why uniform error responses are a design defect rather than a style choice. If every failure returns "Operation failed", the agent cannot distinguish a two-second network blip from an invalid order ID from a refund the policy forbids. Faced with an undifferentiated failure, the model does the only thing available: it retries, or it apologizes and stops. Both are wrong most of the time.

Four categories, one retry decision

The guide separates failures into four kinds, and the practical purpose of the taxonomy is to answer one question — should the agent call this tool again?

Category What went wrong Retryable? What the agent does next
Transient Timeout, service unavailable Yes Retry — the same call may well succeed
Validation Invalid input Not as-is Fix the arguments, then call again
Business Policy violation, such as a refund outside the return window No Relay the customer-friendly explanation; the outcome is a decision, not a fault
Permission The caller lacks rights Not by the agent Take a different path, typically escalation

Encode this explicitly: return errorCategory (transient / validation / permission), an isRetryable boolean, and a human-readable description. For business rule violations add retriable: false together with a customer-friendly explanation the agent can relay verbatim, so it explains the policy instead of inventing one or looping on retries.3

One placement detail the guide does not spell out, and it is the difference between a lesson and a working server. This metadata has to travel inside the result's content (or structuredContent, mirrored into a text block), not alongside it. The model reads the content blocks. A field sitting next to content never reaches it, and the agent sees the same undifferentiated failure you were trying to eliminate. The exam tests which fields to return; your server has to get where right too.

A wording note: the guide is inconsistent about the spelling of this flag — isRetryable where it describes the general error metadata, retriable where it describes business rule violations. Recognize either on the exam, carry one of them in a given payload rather than both, and treat the retryability concept as the thing being tested.

Local recovery, then honest propagation

In a multi-agent system, subagents should absorb transient failures locally — retry within the subagent rather than surfacing every blip to the coordinator.23 Only errors that cannot be resolved locally go up, and when they do they must travel with the full structured error context: failure type, what was attempted, partial results, and alternative approaches. That last element is the one authors forget — naming the alternatives (proceed with reduced coverage, re-delegate the failed queries later) is what lets the coordinator reroute or degrade gracefully instead of restarting the work.

Empty is not broken

Finally, distinguish an access failure (which needs a retry decision) from a valid empty result (a successful query with no matches). lookup_order finding no orders for a customer is a success with an empty list, not an isError. Conflating the two makes agents retry queries that already answered correctly.

Error category drives the agent's recovery actionLet the learner read off, for each of the four MCP error categories, whether the call is retryable and what the agent should do next. Draw it as a 4-row by 2-column grid: the four categories are the rows, the two columns are isRetryable and the agent recovery action, and the relation labels are the cell values. The point is that the category is what makes recovery mechanical.isRetryableAgent recovery actiontransienttransientvalidationvalidationpermissionpermissionbusiness rulebusiness ruletrue — timeout or service unavailabletrueretry with backoff, same call may succeedretry with backoff, same call maysucceedfalse — invalid inputfalsecorrect the arguments, then call againcorrect the arguments, then callagainfalse — caller lacks rightsfalseescalate_to_human, different path neededescalate_to_human, different pathneededfalse — policy violationfalserelay customerExplanation, do not retryrelay customerExplanation, do notretrytimeout or service unavailableinvalid inputcaller lacks rightspolicy violation
Error category drives the agent's recovery action

Let the learner read off, for each of the four MCP error categories, whether the call is retryable and what the agent should do next. Draw it as a 4-row by 2-column grid: the four categories are the rows, the two columns are isRetryable and the agent recovery action, and the relation labels are the cell values. The point is that the category is what makes recovery mechanical.

Local recovery in a subagent, propagation only when neededShow the decision path from an MCP tool result to either local recovery or propagation, and show that a valid empty result bypasses error handling entirely.Subagent calls MCP toolSubagentcalls MCPtoolValid empty resultValid emptyresultisError returnedisErrorreturnedInspect errorCategoryInspecterrorCategoryRetry locallyRetrylocallyResolved locallyResolvedlocallyPropagate to coordinatorPropagatetocoordinatorPartial results plus attemptsPartialresults plusattemptsCoordinator decides next stepCoordinatordecides nextstepquery succeeded,no matchestool failedread structuredmetadatatransient andisRetryable truevalidation,permission,businesssucceededretries exhaustedinclude what wasattemptedreroute, degrade orescalateno retry needed
Local recovery in a subagent, propagation only when needed

Show the decision path from an MCP tool result to either local recovery or propagation, and show that a valid empty result bypasses error handling entirely.

one transition at a time

process_refund declines on policy, not on failure

Scenario 1 · Customer Support Resolution Agent

A customer asks for a refund 90 days after purchase; the return window is 30 days. If process_refund returns isError with "Refund failed", the agent will typically retry two or three times and then escalate a case that never needed a human.

The correct response marks the outcome as a business error with retriable: false and includes a customer-friendly explanation. The agent then does the right thing on the first pass: it tells the customer why the refund cannot be processed and offers the next step, which supports the 80%+ first-contact resolution target instead of burning an escalation.

json
{
  "isError": true,
  "structuredContent": {
    "errorCategory": "business",
    "retriable": false,
    "code": "REFUND_WINDOW_EXPIRED",
    "customerExplanation": "This order was delivered on 12 March, which is outside our 30-day refund window. I can offer store credit or connect you with a specialist to review an exception.",
    "suggestedNextAction": "offer_store_credit_or_escalate"
  },
  "content": [{
    "type": "text",
    "text": "{\"errorCategory\":\"business\",\"retriable\":false,\"code\":\"REFUND_WINDOW_EXPIRED\",\"customerExplanation\":\"This order was delivered on 12 March, which is outside our 30-day refund window. I can offer store credit or connect you with a specialist to review an exception.\",\"suggestedNextAction\":\"offer_store_credit_or_escalate\"}"
  }]
}
MCP tool result: business rule violation

Transient versus validation on the same tool

Scenario 1 · Customer Support Resolution Agent

lookup_order can fail in three genuinely different ways, and the agent's correct behavior differs in each. Returning the same shape with different metadata is what makes the difference actionable: retry the timeout, fix the malformed ID, and treat "no orders" as an answer rather than a fault.

json
// 1. Transient — retry is worthwhile
{ "isError": true,
  "structuredContent": { "errorCategory": "transient", "isRetryable": true,
    "message": "Order service timed out after 5s.", "retryAfterMs": 1000 },
  "content": [{ "type": "text",
    "text": "{\"errorCategory\":\"transient\",\"isRetryable\":true,\"message\":\"Order service timed out after 5s.\",\"retryAfterMs\":1000}" }] }

// 2. Validation — retrying the same call is wasted; fix the argument
{ "isError": true,
  "structuredContent": { "errorCategory": "validation", "isRetryable": false,
    "message": "order_id must match ORD-<6 digits>; received 'last one'.",
    "field": "order_id" },
  "content": [{ "type": "text",
    "text": "{\"errorCategory\":\"validation\",\"isRetryable\":false,\"message\":\"order_id must match ORD-<6 digits>; received 'last one'.\",\"field\":\"order_id\"}" }] }

// 3. Valid empty result — NOT an error
{ "isError": false,
  "structuredContent": { "orders": [], "message": "Customer CUS-4471 has no orders." },
  "content": [{ "type": "text",
    "text": "{\"orders\":[],\"message\":\"Customer CUS-4471 has no orders.\"}" }] }
Three outcomes from one tool, distinguished by metadata

Subagent absorbs the blip, escalates the rest

Scenario 3 · Multi-Agent Research System

The web-search subagent hits a rate limit on its third of eight queries. It should retry locally with backoff rather than telling the coordinator the research failed — the coordinator has no more information than the subagent does, and a round trip costs context on both sides.

If the failure survives local recovery, the subagent propagates it upward with the partial results it did gather, a record of what it attempted, and the alternative approaches it can see. The coordinator can then decide to proceed with six of eight sources, re-delegate the remaining two, or flag reduced coverage in the final report — decisions it cannot make from a bare "search failed". Naming those options in an alternatives field is what turns the report from a complaint into a decision brief.

json
{
  "structuredContent": {
    "status": "partial",
    "errorCategory": "transient",
    "isRetryable": true,
    "message": "Search provider rate-limited queries 3 and 7 after 3 local retries.",
    "attempted": [
      "3 retries with exponential backoff (1s, 2s, 4s)",
      "fallback to cached results — none available"
    ],
    "partialResults": {
      "completedQueries": 6,
      "totalQueries": 8,
      "sources": ["src_101", "src_102", "src_105", "src_107", "src_110", "src_114"]
    },
    "alternatives": [
      "proceed with the 6 completed queries and flag reduced coverage",
      "re-delegate queries 3 and 7 after the rate-limit window resets",
      "substitute the document-analysis subagent for the two missing topics"
    ]
  },
  "content": [{
    "type": "text",
    "text": "{\"status\":\"partial\",\"errorCategory\":\"transient\",\"isRetryable\":true,\"message\":\"Search provider rate-limited queries 3 and 7 after 3 local retries.\",\"attempted\":[\"3 retries with exponential backoff (1s, 2s, 4s)\",\"fallback to cached results — none available\"],\"partialResults\":{\"completedQueries\":6,\"totalQueries\":8,\"sources\":[\"src_101\",\"src_102\",\"src_105\",\"src_107\",\"src_110\",\"src_114\"]},\"alternatives\":[\"proceed with the 6 completed queries and flag reduced coverage\",\"re-delegate queries 3 and 7 after the rate-limit window resets\",\"substitute the document-analysis subagent for the two missing topics\"]}"
  }]
}
Subagent → coordinator report after exhausting local recovery
  • Returning a uniform "Operation failed" for every tool failure instead of structured metadata because the agent cannot make an appropriate recovery decision from an undifferentiated error.
  • Marking a policy violation as a generic retryable error instead of business with retriable: false and a customer-friendly explanation because the agent will burn retries and then escalate a case it could have resolved by explaining the policy.
  • Propagating every transient subagent failure to the coordinator instead of recovering locally first because it wastes coordinator context on blips the subagent could have absorbed.
  • Returning isError for a query that legitimately found nothing instead of a valid empty result because the agent then retries a call that already answered correctly.
  • When a stem shows an agent retrying a doomed operation or apologizing instead of explaining, the credited fix is structured error metadata — errorCategory plus isRetryable / retriable — not more retry logic or a longer system prompt.
  • Distractors often propose a single global retry policy (for example "retry all failures three times"). The guide's position is that the retry decision belongs in the error payload, per category.
  • Multi-agent stems test the propagation rule: recover transient failures locally, and escalate only what cannot be resolved locally — always with partial results and what was attempted. An option that escalates everything, or one that swallows the failure silently, is wrong.
References — 3 sources
  1. Handle tool calls Anthropic The exact `tool_result` message contract, including the case where text placed before `tool_result` blocks ends the turn early and returns a 400.
  2. Tools Model Context Protocol The normative source for `isError`. It splits failures into protocol errors (JSON-RPC) and tool execution errors (`isError: true`), and states the reason this unit's whole argument depends on: "Clients SHOULD provide tool execution errors to language models to enable self-correction." It also fixes the result shape — model-visible data travels in `content` or `structuredContent`, and a tool returning structured content "SHOULD also return the serialized JSON in a TextContent block".
  3. Connect to external tools with MCP Anthropic The SDK's error-handling section, including per-server status reporting (`mcpServerStatus()` / `get_mcp_status()`) — the difference between a tool that failed and a server that never connected, which the four error categories do not cover.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • The MCP isError flag pattern for communicating tool failures back to the agent
  • The distinction between transient errors (timeouts, service unavailability), validation errors (invalid input), business errors (policy violations), and permission errors
  • Why uniform error responses (generic "Operation failed") prevent the agent from making appropriate recovery decisions
  • The difference between retryable and non-retryable errors, and how returning structured metadata prevents wasted retry attempts

Skills in

  • Returning structured error metadata including errorCategory (transient/validation/permission), isRetryable boolean, and human-readable descriptions
  • Including retriable: false flags and customer-friendly explanations for business rule violations so the agent can communicate appropriately
  • Implementing local error recovery within subagents for transient failures, propagating to the coordinator only errors that cannot be resolved locally along with partial results and what was attempted
  • Distinguishing between access failures (needing retry decisions) and valid empty results (representing successful queries with no matches)
Back to top