Note on selection
Enforce structured output using tool use and JSON schemas
What you need to know
- A tool whose
input_schemais your output schema — not a prompt asking for JSON — is what makes output schema-compliant. -
"auto"can return text instead of calling anything; never use it when structured output is mandatory. -
strictmode eliminates syntax errors only, and the live API makes it opt-in. Shape is not correctness. - A required field the source lacks forces a fabricated value; optional or nullable turns a hallucination into an honest null.
- Closed enums break on real data: add
"unclear"and"other"plus a detail string.
Tool use is the schema enforcement mechanism
The reliable way to get structured output is not to ask for JSON in the prompt. Instead you define a tool whose input_schema is your output schema, let the model call it, and read the structured data out of the tool_use block.47 Because the model emits tool input rather than free text, you get schema-compliant output and JSON syntax errors are eliminated: no markdown fences, no preamble, no trailing comma to repair. That is the guide's claim and the exam's answer. In the live API the guarantee is opt-in: you set strict: true on the tool definition and give the schema additionalProperties: false.48 The API also ships a first-class path the guide does not cover: output_config.format returns validated JSON with no extraction tool. Read the reliable way as the way the exam tests. The extraction "tool" usually executes nothing on your side; its schema exists purely to constrain the shape of the answer.
This is the API-layer mechanism. When the surface is Claude Code in CI rather than a direct API call, the equivalent controls are --output-format json plus --json-schema (3.6). Same goal of a schema-enforced machine-readable result, expressed as CLI flags instead of a tool definition.
The three tool_choice modes
The distinction the exam tests:
| Mode | Guarantee | Who picks the tool | Use it when |
|---|---|---|---|
"auto" |
May call a tool, may return plain text | Model | Never, when structured output is mandatory |
"any" |
Must call a tool | Model | Several schemas, document type unknown — the tool name doubles as the classification |
{"type": "tool", "name": "..."} |
Must call that tool | You | One particular extraction has to run, e.g. metadata before enrichment consumes it |
The API has a fourth value, none, which blocks tool calls for that request.19 The guide does not test this — if it appears as an option, it is not the credited answer.
Strict schemas stop syntax errors, not semantic ones
This is the highest-value nuance in the statement. JSON Schema's strict mode is the feature name to know here. The guide lists it as strict mode for syntax error elimination — syntax errors, and nothing more. Note where it actually lives: strict is a field on the tool definition, not on the schema and not on tool_choice.48 Without it, the schema guides the model rather than binding it. A strict schema guarantees the shape: required keys present, values of the right type, enum values legal. It guarantees nothing about meaning. Line items can still fail to sum to the stated total; a vendor tax ID can land in the invoice-number field; a date can be wrong in perfect ISO 8601. Schema compliance is not correctness, and the remedy is semantic validation (4.4), not a stricter schema.
Schema design that does not induce fabrication
Two design decisions carry most of the weight:
- Required vs optional. If a source document may legitimately not contain a field, model it as optional or nullable. A field marked required forces the model to produce something, and what it produces is a fabricated value. Making absence representable is what converts a hallucination into an honest null.49
- Extensible enums. Closed enums break on real-world data. Add an
"unclear"value for genuinely ambiguous cases, and an"other"value paired with a detail string so an unanticipated category can be captured rather than forced into the nearest wrong bucket.
Finally, a schema constrains structure but says nothing about formatting. Source documents write dates, currencies and units inconsistently. Put format normalization rules in the prompt alongside the strict schema: "dates as YYYY-MM-DD; amounts as decimals with no currency symbol; for a range, record the lower bound and note it".
Give the three tool_choice modes as rows against the three columns the exam tests — what the call guarantees, who selects the schema, and which stem it fits — so a stem maps to a mode in one lookup.
Show the concrete call path from document to validated record, and make clear that a strict tool schema removes syntax errors while only the semantic validator can catch wrong values.
Worked examples
An extraction tool schema designed against fabrication
Scenario 6 · Structured Data ExtractionNote what is not required: any field the document may legitimately omit is nullable, so "absent" has a legal representation. The currency enum ships with other plus a detail string and an unclear value, so an unanticipated or genuinely ambiguous currency does not get forced into USD.
{
"name": "extract_invoice",
"description": "Record the fields present in this invoice. Omit or null any field the document does not contain; never infer a value.",
"input_schema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"issue_date": {
"type": ["string", "null"],
"description": "YYYY-MM-DD, or null if the document has no issue date"
},
"vendor_tax_id": { "type": ["string", "null"] },
"currency": {
"type": "string",
"enum": ["USD", "EUR", "GBP", "other", "unclear"]
},
"currency_detail": {
"type": ["string", "null"],
"description": "Required when currency is 'other'; the code as printed"
},
"line_items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": { "type": "string" },
"amount": { "type": "number" }
},
"required": ["description", "amount"]
}
}
},
"required": ["invoice_number", "currency", "line_items"]
}
}Choosing tool_choice: unknown document type vs mandatory step
Scenario 6 · Structured Data ExtractionTwo different problems, two different modes. When the incoming document could be an invoice, a purchase order or a receipt, "any" guarantees structured output and tells you which schema matched via the returned tool name. When a downstream enrichment step requires metadata to already exist, force that specific tool by name. Leaving either case on "auto" risks a plain-text answer that breaks the pipeline.
// Document type unknown: the model must call one of the schemas,
// and its choice of tool name is the classification.
const res = await client.messages.create({
model: 'claude-opus-5',
max_tokens: 4096,
tools: [extractInvoice, extractPurchaseOrder, extractReceipt],
tool_choice: { type: 'any' },
messages: [{ role: 'user', content: [documentBlock] }],
});
const call = res.content.find((b) => b.type === 'tool_use');
// call.name -> which document type matched
// call.input -> the structured extraction, already schema-compliant
// Metadata must exist before the enrichment pass runs: force that one tool.
const meta = await client.messages.create({
model: 'claude-opus-5',
max_tokens: 2048,
tools: [extractMetadata, extractInvoice],
tool_choice: { type: 'tool', name: 'extract_metadata' },
messages: [{ role: 'user', content: [documentBlock] }],
});Normalization rules in the prompt beside the strict schema
Scenario 6 · Structured Data ExtractionThe schema says issue_date is a string and amount is a number. It cannot say that "11/02/26" is ambiguous, that "1.240,00" is European decimal notation, or that "about a tablespoon" is not a number. Those rules belong in the prompt, and they are what stops the model from silently guessing a convention.
Normalization rules:
- Dates: emit YYYY-MM-DD. If the source order is ambiguous (11/02/26) and the
document gives no locale signal, set the date to null and set
date_ambiguous to true. Do not pick an interpretation.
- Amounts: emit a decimal number with no currency symbol and no thousands
separator. Treat "1.240,00" as 1240.00 only when the document uses comma
decimals elsewhere; otherwise set the amount to null.
- Informal quantities ("about a tablespoon", "roughly room temperature"):
leave the numeric field null and copy the phrase verbatim into the
matching *_note field.
- Never convert, round, or infer a unit that the document does not state.Anti-patterns
- Prose JSON requests: instructing the model to "respond with only valid JSON matching this shape" instead of defining a tool with a JSON schema because prose leaves syntax errors, markdown fences and preamble on the table while tool use eliminates them.
- Leaving tool_choice on "auto" when structured output is mandatory instead of using "any" or a forced named tool because with "auto" the model may legitimately answer in plain text and the pipeline breaks.
- All-required schemas: marking every field required so downstream code never sees a null, instead of making genuinely optional fields nullable because the model then fabricates a value to satisfy the schema.
- Treating schema compliance as correctness: skipping semantic checks because the tool-use output validated, instead of verifying sums and field placement because a strict schema eliminates syntax errors and nothing else.
How it is examined
- Map the stem to the tool_choice mode: "document type is unknown, several extraction schemas exist" → "any"; "this extraction must run before the enrichment step" → forced {"type":"tool","name":...}; "structured output is required" → anything but "auto".
- Distractors here are prompt-side or post-processing fixes: "add 'return only JSON' to the prompt", "add a JSON repair/retry parser", "lower output length". The credited answer is tool use with a JSON schema.
- If the stem describes valid JSON whose values are wrong (line items that do not sum, a value in the wrong field), the fault is semantic and the answer lives in 4.4 — not in a stricter schema or more required fields.
Beyond the exam — what the API does that this does not grade
strict: true is a top-level field on the tool definition, beside name and input_schema. It is not a schema keyword and not a tool_choice option. The schema must declare additionalProperties: false and required.48
Strict mode also drops several JSON Schema keywords: recursive $ref, numeric constraints (minimum, maximum, multipleOf) and string constraints (minLength, maxLength). Validate those in your own code. Schemas written for plain tool use routinely carry them — the CI schema in 3.6 uses "minimum": 1 — and under strict mode they stop being enforced.
There is also a second mechanism the guide does not cover. output_config: {format: {type: "json_schema", schema: ...}} constrains the response rather than a tool input, for the case where no tool call is wanted. It is incompatible with citations.47
References — 4 sources
- Define tools Anthropic The normative tool-definition reference — name, `description`, `input_schema` — with Anthropic's own guidance on description quality, and the place to check whether the unit's four description components are the source's four. It is also where `tool_choice` is enumerated: "there are four possible options", `auto`, `any`, `tool` and `none`. With `any` or `tool` the API prefills the assistant message, so no natural-language text precedes the `tool_use` block.
- Structured outputs Anthropic Documents a first-class structured-output path — `output_config.format`, formerly the `output_format` beta — that returns validated JSON directly without the extraction-tool trick, and explains how it differs from strict tool use. It also carries the supported Pydantic integration (`client.messages.parse()` with `output_format=Model`), which shows where schema validation stops and 4.4's semantic validator must begin.
- Strict tool use Anthropic The actual feature: `strict: true` on a tool definition, enforced by grammar-constrained sampling. It requires `additionalProperties: false` and `required`, and it drops several JSON Schema keywords — recursive `$ref`, `minimum`, `maximum`, `multipleOf`, `minLength`, `maxLength`.
- Reduce hallucinations Anthropic Anthropic's guidance on giving the model a legal way to say "not present", which is the general principle the nullable-field rule implements at the schema layer.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Tool use (tool_use) with JSON schemas as the most reliable approach for guaranteed schema- compliant structured output, eliminating JSON syntax errors
- The distinction between tool_choice: "auto" (model may return text instead of calling a tool), "any" (model must call a tool but can choose which), and forced tool selection (model must call a specific named tool)
- That strict JSON schemas via tool use eliminate syntax errors but do not prevent semantic errors (e.g., line items that don't sum to total, values in wrong fields)
- Schema design considerations: required vs optional fields, enum fields with "other" + detail string patterns for extensible categories
Skills in
- Defining extraction tools with JSON schemas as input parameters and extracting structured data from the tool_use response
- Setting tool_choice: "any" to guarantee structured output when multiple extraction schemas exist and the document type is unknown
- Forcing a specific tool with tool_choice: {"type": "tool", "name": "extract_metadata"} to ensure a particular extraction runs before enrichment steps
- Designing schema fields as optional (nullable) when source documents may not contain the information, preventing the model from fabricating values to satisfy required fields
- Adding enum values like "unclear" for ambiguous cases and "other" + detail fields for extensible categorization
- Including format normalization rules in prompts alongside strict output schemas to handle inconsistent source formatting