Note on selection
Implement structured error responses for MCP tools
What you need to know
- MCP signals failure with the isError flag, but the flag alone carries no recovery information — the structured metadata you attach does.
- Return errorCategory (transient / validation / permission), an isRetryable boolean, and a human-readable description on every failure.
- Business rule violations get retriable: false plus a customer-friendly explanation the agent can relay, so it stops retrying and communicates the policy. The guide spells the flag both isRetryable and retriable — recognize either.
- Uniform "Operation failed" responses force the agent to guess, which produces wasted retries or premature give-up.
- Subagents recover from transient failures locally and propagate to the coordinator only what they cannot resolve — with the four elements of structured error context: failure type, what was attempted, partial results, and alternative approaches.
- A successful query with no matches is a valid empty result, not an error; conflating it with an access failure triggers pointless retries.
An error is a message to a reasoning system
MCP communicates tool failure back to the agent with the isError flag on the tool result. That flag says only "this call failed" — it carries no recovery information by itself. Everything the agent needs in order to choose a next action has to be in the payload you attach to it.22
This is why uniform error responses are a design defect rather than a style choice. If every failure returns "Operation failed", the agent cannot distinguish a two-second network blip from an invalid order ID from a refund the policy forbids. Faced with an undifferentiated failure, the model does the only thing available: it retries, or it apologizes and stops. Both are wrong most of the time.
Four categories, one retry decision
The guide separates failures into four kinds, and the practical purpose of the taxonomy is to answer one question — should the agent call this tool again?
| Category | What went wrong | Retryable? | What the agent does next |
|---|---|---|---|
| Transient | Timeout, service unavailable | Yes | Retry — the same call may well succeed |
| Validation | Invalid input | Not as-is | Fix the arguments, then call again |
| Business | Policy violation, such as a refund outside the return window | No | Relay the customer-friendly explanation; the outcome is a decision, not a fault |
| Permission | The caller lacks rights | Not by the agent | Take a different path, typically escalation |
Encode this explicitly: return errorCategory (transient / validation / permission), an isRetryable boolean, and a human-readable description. For business rule violations add retriable: false together with a customer-friendly explanation the agent can relay verbatim, so it explains the policy instead of inventing one or looping on retries.3
One placement detail the guide does not spell out, and it is the difference between a lesson and a working server. This metadata has to travel inside the result's content (or structuredContent, mirrored into a text block), not alongside it. The model reads the content blocks. A field sitting next to content never reaches it, and the agent sees the same undifferentiated failure you were trying to eliminate. The exam tests which fields to return; your server has to get where right too.
A wording note: the guide is inconsistent about the spelling of this flag — isRetryable where it describes the general error metadata, retriable where it describes business rule violations. Recognize either on the exam, carry one of them in a given payload rather than both, and treat the retryability concept as the thing being tested.
Local recovery, then honest propagation
In a multi-agent system, subagents should absorb transient failures locally — retry within the subagent rather than surfacing every blip to the coordinator.23 Only errors that cannot be resolved locally go up, and when they do they must travel with the full structured error context: failure type, what was attempted, partial results, and alternative approaches. That last element is the one authors forget — naming the alternatives (proceed with reduced coverage, re-delegate the failed queries later) is what lets the coordinator reroute or degrade gracefully instead of restarting the work.
Empty is not broken
Finally, distinguish an access failure (which needs a retry decision) from a valid empty result (a successful query with no matches). lookup_order finding no orders for a customer is a success with an empty list, not an isError. Conflating the two makes agents retry queries that already answered correctly.
Let the learner read off, for each of the four MCP error categories, whether the call is retryable and what the agent should do next. Draw it as a 4-row by 2-column grid: the four categories are the rows, the two columns are isRetryable and the agent recovery action, and the relation labels are the cell values. The point is that the category is what makes recovery mechanical.
Show the decision path from an MCP tool result to either local recovery or propagation, and show that a valid empty result bypasses error handling entirely.
Worked examples
process_refund declines on policy, not on failure
Scenario 1 · Customer Support Resolution AgentA customer asks for a refund 90 days after purchase; the return window is 30 days. If process_refund returns isError with "Refund failed", the agent will typically retry two or three times and then escalate a case that never needed a human.
The correct response marks the outcome as a business error with retriable: false and includes a customer-friendly explanation. The agent then does the right thing on the first pass: it tells the customer why the refund cannot be processed and offers the next step, which supports the 80%+ first-contact resolution target instead of burning an escalation.
{
"isError": true,
"structuredContent": {
"errorCategory": "business",
"retriable": false,
"code": "REFUND_WINDOW_EXPIRED",
"customerExplanation": "This order was delivered on 12 March, which is outside our 30-day refund window. I can offer store credit or connect you with a specialist to review an exception.",
"suggestedNextAction": "offer_store_credit_or_escalate"
},
"content": [{
"type": "text",
"text": "{\"errorCategory\":\"business\",\"retriable\":false,\"code\":\"REFUND_WINDOW_EXPIRED\",\"customerExplanation\":\"This order was delivered on 12 March, which is outside our 30-day refund window. I can offer store credit or connect you with a specialist to review an exception.\",\"suggestedNextAction\":\"offer_store_credit_or_escalate\"}"
}]
}Transient versus validation on the same tool
Scenario 1 · Customer Support Resolution Agentlookup_order can fail in three genuinely different ways, and the agent's correct behavior differs in each. Returning the same shape with different metadata is what makes the difference actionable: retry the timeout, fix the malformed ID, and treat "no orders" as an answer rather than a fault.
// 1. Transient — retry is worthwhile
{ "isError": true,
"structuredContent": { "errorCategory": "transient", "isRetryable": true,
"message": "Order service timed out after 5s.", "retryAfterMs": 1000 },
"content": [{ "type": "text",
"text": "{\"errorCategory\":\"transient\",\"isRetryable\":true,\"message\":\"Order service timed out after 5s.\",\"retryAfterMs\":1000}" }] }
// 2. Validation — retrying the same call is wasted; fix the argument
{ "isError": true,
"structuredContent": { "errorCategory": "validation", "isRetryable": false,
"message": "order_id must match ORD-<6 digits>; received 'last one'.",
"field": "order_id" },
"content": [{ "type": "text",
"text": "{\"errorCategory\":\"validation\",\"isRetryable\":false,\"message\":\"order_id must match ORD-<6 digits>; received 'last one'.\",\"field\":\"order_id\"}" }] }
// 3. Valid empty result — NOT an error
{ "isError": false,
"structuredContent": { "orders": [], "message": "Customer CUS-4471 has no orders." },
"content": [{ "type": "text",
"text": "{\"orders\":[],\"message\":\"Customer CUS-4471 has no orders.\"}" }] }Subagent absorbs the blip, escalates the rest
Scenario 3 · Multi-Agent Research SystemThe web-search subagent hits a rate limit on its third of eight queries. It should retry locally with backoff rather than telling the coordinator the research failed — the coordinator has no more information than the subagent does, and a round trip costs context on both sides.
If the failure survives local recovery, the subagent propagates it upward with the partial results it did gather, a record of what it attempted, and the alternative approaches it can see. The coordinator can then decide to proceed with six of eight sources, re-delegate the remaining two, or flag reduced coverage in the final report — decisions it cannot make from a bare "search failed". Naming those options in an alternatives field is what turns the report from a complaint into a decision brief.
{
"structuredContent": {
"status": "partial",
"errorCategory": "transient",
"isRetryable": true,
"message": "Search provider rate-limited queries 3 and 7 after 3 local retries.",
"attempted": [
"3 retries with exponential backoff (1s, 2s, 4s)",
"fallback to cached results — none available"
],
"partialResults": {
"completedQueries": 6,
"totalQueries": 8,
"sources": ["src_101", "src_102", "src_105", "src_107", "src_110", "src_114"]
},
"alternatives": [
"proceed with the 6 completed queries and flag reduced coverage",
"re-delegate queries 3 and 7 after the rate-limit window resets",
"substitute the document-analysis subagent for the two missing topics"
]
},
"content": [{
"type": "text",
"text": "{\"status\":\"partial\",\"errorCategory\":\"transient\",\"isRetryable\":true,\"message\":\"Search provider rate-limited queries 3 and 7 after 3 local retries.\",\"attempted\":[\"3 retries with exponential backoff (1s, 2s, 4s)\",\"fallback to cached results — none available\"],\"partialResults\":{\"completedQueries\":6,\"totalQueries\":8,\"sources\":[\"src_101\",\"src_102\",\"src_105\",\"src_107\",\"src_110\",\"src_114\"]},\"alternatives\":[\"proceed with the 6 completed queries and flag reduced coverage\",\"re-delegate queries 3 and 7 after the rate-limit window resets\",\"substitute the document-analysis subagent for the two missing topics\"]}"
}]
}Anti-patterns
- Returning a uniform "Operation failed" for every tool failure instead of structured metadata because the agent cannot make an appropriate recovery decision from an undifferentiated error.
- Marking a policy violation as a generic retryable error instead of business with retriable: false and a customer-friendly explanation because the agent will burn retries and then escalate a case it could have resolved by explaining the policy.
- Propagating every transient subagent failure to the coordinator instead of recovering locally first because it wastes coordinator context on blips the subagent could have absorbed.
- Returning isError for a query that legitimately found nothing instead of a valid empty result because the agent then retries a call that already answered correctly.
How it is examined
- When a stem shows an agent retrying a doomed operation or apologizing instead of explaining, the credited fix is structured error metadata — errorCategory plus isRetryable / retriable — not more retry logic or a longer system prompt.
- Distractors often propose a single global retry policy (for example "retry all failures three times"). The guide's position is that the retry decision belongs in the error payload, per category.
- Multi-agent stems test the propagation rule: recover transient failures locally, and escalate only what cannot be resolved locally — always with partial results and what was attempted. An option that escalates everything, or one that swallows the failure silently, is wrong.
References — 3 sources
- Handle tool calls Anthropic The exact `tool_result` message contract, including the case where text placed before `tool_result` blocks ends the turn early and returns a 400.
- Tools Model Context Protocol The normative source for `isError`. It splits failures into protocol errors (JSON-RPC) and tool execution errors (`isError: true`), and states the reason this unit's whole argument depends on: "Clients SHOULD provide tool execution errors to language models to enable self-correction." It also fixes the result shape — model-visible data travels in `content` or `structuredContent`, and a tool returning structured content "SHOULD also return the serialized JSON in a TextContent block".
- Connect to external tools with MCP Anthropic The SDK's error-handling section, including per-server status reporting (`mcpServerStatus()` / `get_mcp_status()`) — the difference between a tool that failed and a server that never connected, which the four error categories do not cover.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- The MCP isError flag pattern for communicating tool failures back to the agent
- The distinction between transient errors (timeouts, service unavailability), validation errors (invalid input), business errors (policy violations), and permission errors
- Why uniform error responses (generic "Operation failed") prevent the agent from making appropriate recovery decisions
- The difference between retryable and non-retryable errors, and how returning structured metadata prevents wasted retry attempts
Skills in
- Returning structured error metadata including errorCategory (transient/validation/permission), isRetryable boolean, and human-readable descriptions
- Including retriable: false flags and customer-friendly explanations for business rule violations so the agent can communicate appropriately
- Implementing local error recovery within subagents for transient failures, propagating to the coordinator only errors that cannot be resolved locally along with partial results and what was attempted
- Distinguishing between access failures (needing retry decisions) and valid empty results (representing successful queries with no matches)