Note on selection
Design effective escalation and ambiguity resolution patterns
What you need to know
- Valid triggers: an explicit request for a human, a policy exception or policy gap, and inability to make meaningful progress — not "the case looks complex".
- An explicit demand for a human is honored immediately, with no investigation attempt first.
- Frustration without an explicit request: acknowledge it, offer to resolve if the issue is in scope, and escalate only if the customer reiterates their preference.
- Sentiment tracks the outcome, not case difficulty, and self-reported confidence is least calibrated on exactly the novel hard cases escalation exists for — unlike calibrated, field-level confidence, which legitimately prioritizes offline human review (5.5) or routes an already-generated finding (4.6).
- Multiple tool matches mean ask for another identifier — never pick by heuristic.
- Escalation criteria in the system prompt need few-shot examples of both escalating and resolving to be usable.
Escalation is a design problem, not a fallback.14 The exam tests whether you can state which signals justify handing off to a human and which ones only look like they do.
The signals, evaluated in this order. The order carries as much weight as the list. An explicit request for a person short-circuits every check below it, and no lower signal overrides one above.
| Check, in this order | What fires it | What the agent does | Why not the alternative |
|---|---|---|---|
| 1 | The customer explicitly asks for a person | Call escalate_to_human now, with no investigation first |
Working the case after a direct request reads as refusing it |
| 2 | A policy exception, or the policy is silent on the request | Escalate as a policy gap, stating what the policy does cover | Any autonomous answer invents an entitlement, or invents a denial |
| 3 | No meaningful progress left with the available tools | Escalate, carrying what was already attempted | Another lap through the same tools returns the same non-answer |
| 4 | get_customer returns multiple matches |
Ask for one more identifier — order number, postal code, last four card digits — then re-run | Picking by "most recent activity" or "closest name" acts on the wrong person |
| 5 | Frustration, but no explicit request for a person | Acknowledge it, then offer to resolve now if the issue is in scope | Escalating on tone abandons a case the agent could have closed |
| Never | The case looks complex, or a self-reported score is low | Keep working | Complexity is what the agent is for |
Row 2 is the one to read twice. The guide's example is competitor price matching when the policy only addresses adjustments on your own site. The policy is silent, so no autonomous answer is correct and inventing one is worse than escalating. Row 5 escalates only if the customer then reiterates a preference for a person: the distinguishing question is whether an explicit request was made, not how negative the message sounds.
This ordering is the guide's own; no published source ranks these triggers. Answer with it as assessed content, not as a general claim about how escalation works. Row 3 has a shipped instance: Claude Code's auto mode escalates after 3 consecutive or 20 total permission denials — counted and externally observable, not a self-report.58 Anthropic's support guidance supplies the missing metric: escalation accuracy, the share of correctly escalated conversations, targeted at 95% or higher.56
Why the two tempting proxies fail. Sentiment-based escalation confuses tone with difficulty. Tone is not noise — escalated support tickets do measurably skew negative.62 But it is a correlate of the outcome, not a measure of whether this agent can resolve this case. Routing on it abandons cases the agent could have closed. Anthropic's own support guidance draws the same line. Sentiment appears there as a business-impact metric measured at the start and end of a conversation, never as a routing signal.56
Self-reported confidence is systematically overconfident, and — decisively for this use — its calibration is weakest on exactly the cases escalation exists for. Models are not simply guessing. Anthropic's own research finds larger models are well calibrated on multiple-choice and true/false questions given the right format, and that P(True) self-evaluation calibrates and improves with scale.59 But that same Anthropic work finds models "struggle with calibration of P(IK) on new tasks". Independent evaluation finds LLMs "tend to be overconfident" when verbalizing a score, with every elicitation method struggling on challenging tasks.60 A live escalation router asks for a confidence number on the novel, hard case where the calibration evidence is weakest. And, unlike the offline setting in 5.5, it has no labeled outcome to calibrate that number against. That is why the score cannot be the trigger: not an absence of introspection, but the absence of calibration where it would have to hold.
Official Sample Question 3 makes the failure precise. An agent that escalates straightforward damage replacements while attempting policy-exception cases autonomously is already incorrectly confident on the hard cases. A "confidence below threshold" router therefore escalates the easy tickets and keeps the hard ones. Both proxies waste human capacity and miss real escalations.
Scope this claim carefully, because it is the axis the exam disambiguates. Self-reported confidence is rejected here as a proxy for case complexity, and as a substitute for the explicit escalation triggers — not as a signal in general. Field-level confidence that has been calibrated against labeled validation sets is legitimate elsewhere: it prioritizes offline human review of extractions (5.5), and it can route a finding that has already been generated (4.6). What it cannot do is decide, live in a conversation, whether this case needs a human. And it must never act as an upstream filter on what counts as a finding at all (4.1).
Why row 4 asks instead of picking. Acting on the wrong customer record is a data-exposure and wrong-refund incident, not a UX inconvenience. A disambiguation heuristic is right often enough to feel safe, and wrong often enough to act on the wrong person. The SDK gives the question a mechanism: AskUserQuestion and the canUseTool callback surface it mid-run.57
How you implement it. Put explicit escalation criteria in the system prompt and pair them with few-shot examples that demonstrate both branches — escalate here, resolve autonomously there. Criteria alone under-specify the boundary; the worked examples are what make the boundary learnable.
Fix the order in which triggers are evaluated: an explicit human request short-circuits everything, policy gaps escalate, multiple matches trigger clarification, and frustration alone leads to an offer to resolve.
Click a decision node to highlight its two outgoing branches and the cue on each label.
Worked examples
System-prompt escalation criteria with both branches
Scenario 1 · Customer Support Resolution AgentThe criteria list makes the triggers explicit, and the few-shot pairs demonstrate the boundary in both directions. Note that the "resolve" example is as important as the "escalate" one: without it the agent over-escalates and first-contact resolution collapses well below the 80% target.
ESCALATE with escalate_to_human when ANY of these hold:
1. The customer explicitly asks to speak with a person.
-> Escalate immediately. Do NOT investigate first.
2. Policy does not cover the request, or an exception is required.
3. You cannot make meaningful progress after your available tools.
Do NOT escalate merely because the case is complex or the customer is upset.
Example A — escalate immediately
Customer: "Stop with the bot answers, transfer me to a human."
Action: escalate_to_human(reason="explicit_human_request")
Example B — resolve, do not escalate
Customer: "This is absurd, the jacket arrived torn and I want it gone."
Action: acknowledge the frustration, confirm the order is inside the
return window, and offer the return now. Escalate only if the customer
then repeats that they want a person.A policy gap, not a hard case
Scenario 1 · Customer Support Resolution AgentA customer asks for a price match against a competitor's listing. The refund policy documents price adjustments for purchases on your own site within 14 days and says nothing about competitors. The agent has enough information to answer — and that is the trap. Because the policy is silent, any autonomous answer either invents an entitlement or denies one that may exist. The correct behavior is to escalate as a policy gap, stating what the policy does cover so the human starts informed.
{
"tool": "escalate_to_human",
"input": {
"reason": "policy_gap",
"customer_request": "price match against competitor listing at USD 79.00",
"policy_coverage": "own-site price adjustments within 14 days only",
"policy_silent_on": "competitor price matching",
"case_facts": {
"order_id": "A-99231",
"order_total": "USD 129.00",
"order_date": "2026-07-22"
},
"agent_recommendation": "no autonomous decision made"
}
}Multiple customer matches: ask, do not disambiguate
Scenario 1 · Customer Support Resolution Agentget_customer on "J. Alvarez" returns three records. A heuristic such as "use the account with the most recent order" is wrong often enough to cause refunds against the wrong account and disclosure of one customer's order history to another. The instruction is to surface the ambiguity to the customer and request an additional identifier that is safe to ask for.
When get_customer returns more than one match:
- Do NOT select a record using recency, name similarity, or order volume.
- Ask the customer for ONE additional identifier: order number,
billing postal code, or the last four digits of the payment card.
- Re-run get_customer with the identifier before any account action.
- If still ambiguous after one clarification round, escalate with the
number of candidate matches and the identifiers already tried.Anti-patterns
- Investigating first when a customer has explicitly asked for a human instead of escalating immediately because continuing to work the case after a direct request reads as refusing the request.
- Using detected sentiment or a self-reported confidence threshold as the live escalation trigger instead of the explicit criteria because tone measures the outcome and self-reported confidence loses its calibration on the novel, hard cases — note this is narrower than "never use confidence": calibrated field-level confidence still prioritizes offline review (5.5).
- Selecting among multiple customer matches with a heuristic instead of asking for an additional identifier because acting on the wrong record causes wrong refunds and cross-customer data exposure.
- Answering a request the policy is silent about instead of escalating the policy gap because an invented entitlement (or an invented denial) is a decision the agent has no basis to make.
How it is examined
- Expect stems that quote a customer utterance. Look for an explicit request for a human: if present, the credited answer escalates immediately and every "investigate first" option is wrong.
- When a stem offers a sentiment classifier or a "confidence below 0.7" threshold as the escalation mechanism, that is the designed distractor — Sample Question 3 credits explicit criteria with few-shot examples instead because the agent is already incorrectly confident on the hard cases. Read the option carefully though: a calibrated field-level score prioritizing offline review (5.5) or routing an existing finding (4.6) is credited, not penalized.
- Policy-gap stems are worded so the agent could plausibly answer (e.g. competitor price matching). The credited answer escalates because the policy is silent, not because the case is hard.
References — 7 sources
- Building effective AI agents Anthropic Engineering, 19 Dec 2024 The source of the vocabulary, with all five named patterns and the workflow-versus-agent distinction that the two-row table here is a projection of.
- Customer support agent Anthropic The only Anthropic page treating escalation as designed behavior. Supplies the missing metric — escalation accuracy, target 95% — and places sentiment among business-impact metrics, never among routing signals.
- Handle approvals and user input Anthropic The SDK mechanism behind "ask for another identifier": `AskUserQuestion` and the `canUseTool` callback. Notes that `AskUserQuestion` is unavailable inside subagents spawned via the Agent tool.
- How we built Claude Code auto mode: a safer way to skip permissions Anthropic Engineering Anthropic's own shipped escalation trigger: three consecutive or twenty total denials stop the model and escalate to a human. A counted, externally observable signal rather than a self-reported score.
- Language Models (Mostly) Know What They Know Kadavath et al., Anthropic (2022) Larger models are well calibrated on multiple-choice and true/false questions given the right format, and `P(True)` self-evaluation improves with scale — but models struggle with calibration of P(IK) on new tasks.
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs Xiong et al., ICLR 2024 LLMs verbalizing confidence "tend to be overconfident, potentially imitating human patterns", and every elicitation method evaluated struggles on challenging tasks.
- How angry are your customers? Sentiment analysis of support tickets that escalate Werner et al., AffectRE @ RE 2018 The one empirical study on the question: a considerable sentiment difference between escalated and non-escalated support tickets. Association with escalation, not with complexity, and the authors call it preliminary.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Appropriate escalation triggers: customer requests for a human, policy exceptions/gaps (not just complex cases), and inability to make meaningful progress
- The distinction between escalating immediately when a customer explicitly demands it versus offering to resolve when the issue is straightforward
- Why sentiment-based escalation and self-reported confidence scores are unreliable proxies for actual case complexity
- How multiple customer matches require clarification (requesting additional identifiers) rather than heuristic selection
Skills in
- Adding explicit escalation criteria with few-shot examples to the system prompt demonstrating when to escalate versus resolve autonomously
- Honoring explicit customer requests for human agents immediately without first attempting investigation
- Acknowledging frustration while offering resolution when the issue is within the agent's capability, escalating only if the customer reiterates their preference
- Escalating when policy is ambiguous or silent on the customer's specific request (e.g., competitor price matching when policy only addresses own-site adjustments)
- Instructing the agent to ask for additional identifiers when tool results return multiple matches, rather than selecting based on heuristics