Note on selection
Design effective tool interfaces with clear descriptions and boundaries
What you need to know
- Tool descriptions are the primary mechanism the model uses to select tools — treat them as production routing logic, not documentation.
- A good description states input format, output contract, example queries, edge cases, and when to use a sibling tool instead.
- Ambiguous or overlapping descriptions (analyze_content vs analyze_document) cause misrouting; fix by renaming and re-scoping, not by adding emphasis.
- A generic tool gives the model no contract to route on. Purpose-specific tools with declared inputs and outputs do.
- Keyword-sensitive system prompt wording can override a well-written tool description — audit the system prompt when routing is wrong.
The description is the routing table
An LLM never reads your implementation — it reads the tool's name, description, and input schema. That text is the routing logic. A one-line description like "Analyzes content" tells the model nothing about boundaries, so as soon as two similar tools are in scope, selection becomes close to arbitrary. Minimal descriptions do not fail loudly; they fail as intermittent misrouting that looks like model unreliability.
A description that routes reliably carries four things beyond a summary of purpose:
- Input format — precisely what the arguments are (a web URL, an order ID, a document ID, a claim string), and their expected shape.
- Output contract — what comes back, so the model can plan the next step instead of guessing.
- Example queries — one or two realistic requests for which this tool is the right answer.
- Boundaries and edge cases — when not to use it, and which sibling tool to use instead.19
Overlap is the canonical failure
The exam's worked failure is analyze_content versus analyze_document with near-identical descriptions. Nothing in either text tells the model which one owns web pages and which owns uploaded files, so calls land on whichever tool the sampled tokens favor. There are two structural fixes, and both are testable skills:
- Rename and re-scope. Rename
analyze_contenttoextract_web_resultsand give it a web-specific description. The name itself now carries a boundary, and the description reinforces it. - Split the generic tool. Break
analyze_documentinto purpose-specific tools with defined input/output contracts:extract_data_points,summarize_content, andverify_claim_against_source. Each has a distinct output shape, which makes the choice mechanical rather than interpretive.
Note the direction of travel: fewer overlapping tools, more purpose-specific ones.20 Splitting is not the same as bloating the tool set — see 2.3 for the count constraint that bounds it.
The system prompt can override a good description
Tool selection is influenced by the whole prompt, not just the tool block. Keyword-sensitive instructions create unintended associations: a system prompt that says "always start by searching" biases the model toward any tool whose name or description contains "search", even when a better-matched tool exists. When a well-described tool is still being bypassed, review the system prompt for keyword-sensitive wording before rewriting the description again.21
Show that misrouting comes from overlapping descriptions, and that the two fixes are renaming with a scoped description and splitting a generic tool into purpose-specific tools with their own contracts. The graph converges top-down rather than splitting into a left and a right half: the two vague tools head it, the renamed and split tools sit below them, and the broken-versus-fixed contrast is carried by color — vague tools and the misrouted call in red, the purpose-specific tools and the correct call in green. The three split tools share one grouped box so the fixed side stays readable.
Worked examples
Two research tools that keep getting confused
Scenario 3 · Multi-Agent Research SystemThe research system's web-search subagent and document-analysis subagent each expose an "analyze" tool. Both descriptions read like generic content analysis, so the coordinator's calls land on the wrong one roughly half the time and the document agent receives URLs it cannot open.
The correct approach is not to add "IMPORTANT: only use for documents" to one description. It is to eliminate the functional overlap: rename the web-side tool to extract_web_results with a description that names web search result pages as its only input, and split the document-side tool into purpose-specific tools. After the change, the input format alone disambiguates the choice.
{
"tools": [
{ "name": "analyze_content", "description": "Analyzes content and returns insights." },
{ "name": "analyze_document", "description": "Analyzes a document and returns insights." }
]
}A description written to be routed on
Scenario 3 · Multi-Agent Research SystemAfter the split, each tool states input format, output contract, an example query, and an explicit boundary that names the alternative. The boundary sentence is what prevents the model from reaching for a neighboring tool when the request is phrased loosely.
{
"name": "verify_claim_against_source",
"description": "Checks whether a single factual claim is supported by one already-loaded source document. Input: 'claim' (one sentence) and 'document_id' from load_document. Output: {supported: boolean, evidence_quote: string, confidence: number}. Example: verify that 'revenue grew 12% in 2024' is supported by document doc_418. Do NOT use this to pull numbers out of a document (use extract_data_points) or to condense it (use summarize_content). Returns supported=false with an empty evidence_quote when the document simply does not discuss the claim.",
"input_schema": {
"type": "object",
"properties": {
"claim": { "type": "string" },
"document_id": { "type": "string" }
},
"required": ["claim", "document_id"]
}
}The system prompt sabotaging a correct tool set
Scenario 1 · Customer Support Resolution AgentThe support agent has four well-described MCP tools: get_customer, lookup_order, process_refund, and escalate_to_human. Descriptions are precise, yet the agent keeps calling lookup_order for pure account questions.
The cause is in the system prompt: "For any customer issue, first look up the relevant order." That keyword-sensitive instruction creates an unintended association between every request and lookup_order, overriding the tool descriptions. The fix is to rewrite the instruction so it is conditional on the request type — "when the request concerns a purchase, shipment, or refund, look up the order first" — rather than weakening or padding the tool descriptions.
Anti-patterns
- Shipping a one-line description ("Analyzes content") instead of stating input format, output contract, examples, and boundaries because minimal descriptions make selection among similar tools unreliable.
- Keeping two near-identically described tools and adding emphatic wording ("ALWAYS use this for documents") instead of renaming and re-scoping them because the overlap itself — not the volume — is what causes misrouting.
- Leaving one generic tool (analyze_document) instead of splitting it into purpose-specific tools with defined input/output contracts because a generic tool gives the model no contract to route on.
- Rewriting tool descriptions again instead of auditing the system prompt for keyword-sensitive instructions because prompt keywords can create tool associations that override even a well-written description.
How it is examined
- Stems typically describe an agent calling the wrong one of two minimally described tools and ask for the most effective first step. The credited answer is description quality: expand each tool's description with the input formats it handles, example queries, edge cases, and boundaries explaining when to use it versus the similar tool. Two distractor families sit either side of it — prompt-side fixes (few-shot examples, emphasis, priority language, a "choose carefully" instruction), which add tokens without addressing the root cause, and jumping straight to restructuring the tool set (consolidating the two tools, or adding a routing layer), which is more effort than a "first step" warrants.
- When the stem says descriptions are already detailed and correct but routing is still wrong, the answer is almost always in the system prompt — look for a keyword-sensitive instruction creating an unintended tool association.
- Watch for distractors that "solve" overlap by consolidating the two tools into one generic tool. The guide calls that "a valid architectural choice" — it is legitimate architecture, just the wrong first response when the immediate problem is inadequate descriptions. Renaming, re-scoping and splitting are the right moves once the overlap is genuinely structural rather than a description gap.
References — 3 sources
- Define tools Anthropic The normative tool-definition reference — name, `description`, `input_schema` — with Anthropic's own guidance on description quality, and the place to check whether the unit's four description components are the source's four. It is also where `tool_choice` is enumerated: "there are four possible options", `auto`, `any`, `tool` and `none`. With `any` or `tool` the API prefills the assistant message, so no natural-language text precedes the `tool_use` block.
- Writing effective tools for agents — with agents Anthropic Engineering The `analyze_content` versus `analyze_document` problem treated at length, with the opposite default: namespacing (`asana_search` vs `jira_search`) and consolidation over splitting. That is the tension this unit resolves by assertion when it says "fewer overlapping, more purpose-specific".
- Troubleshooting tool use Anthropic The documented diagnostic order for wrong-tool-selected symptoms — where a reader checks whether auditing the system prompt for keyword-sensitive wording really is the recommended first move.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Tool descriptions as the primary mechanism LLMs use for tool selection; minimal descriptions lead to unreliable selection among similar tools
- The importance of including input formats, example queries, edge cases, and boundary explanations in tool descriptions
- How ambiguous or overlapping tool descriptions cause misrouting (e.g., analyze_content vs analyze_document with near-identical descriptions)
- The impact of system prompt wording on tool selection: keyword-sensitive instructions can create unintended tool associations
Skills in
- Writing tool descriptions that clearly differentiate each tool's purpose, expected inputs, outputs, and when to use it versus similar alternatives
- Renaming tools and updating descriptions to eliminate functional overlap (e.g., renaming analyze_content to extract_web_results with a web-specific description)
- Splitting generic tools into purpose-specific tools with defined input/output contracts (e.g., splitting a generic analyze_document into extract_data_points, summarize_content, and verify_claim_against_source)
- Reviewing system prompts for keyword-sensitive instructions that might override well-written tool descriptions