Skip to content
CCAR-FAcademy
Domain 4 · Statement 4.3 3 of 6
4.3

Enforce structured output using tool use and JSON schemas

  • A tool whose input_schema is your output schema — not a prompt asking for JSON — is what makes output schema-compliant.
  • "auto" can return text instead of calling anything; never use it when structured output is mandatory.
  • strict mode eliminates syntax errors only, and the live API makes it opt-in. Shape is not correctness.
  • A required field the source lacks forces a fabricated value; optional or nullable turns a hallucination into an honest null.
  • Closed enums break on real data: add "unclear" and "other" plus a detail string.

Tool use is the schema enforcement mechanism

The reliable way to get structured output is not to ask for JSON in the prompt. Instead you define a tool whose input_schema is your output schema, let the model call it, and read the structured data out of the tool_use block.47 Because the model emits tool input rather than free text, you get schema-compliant output and JSON syntax errors are eliminated: no markdown fences, no preamble, no trailing comma to repair. That is the guide's claim and the exam's answer. In the live API the guarantee is opt-in: you set strict: true on the tool definition and give the schema additionalProperties: false.48 The API also ships a first-class path the guide does not cover: output_config.format returns validated JSON with no extraction tool. Read the reliable way as the way the exam tests. The extraction "tool" usually executes nothing on your side; its schema exists purely to constrain the shape of the answer.

This is the API-layer mechanism. When the surface is Claude Code in CI rather than a direct API call, the equivalent controls are --output-format json plus --json-schema (3.6). Same goal of a schema-enforced machine-readable result, expressed as CLI flags instead of a tool definition.

The three tool_choice modes

The distinction the exam tests:

Mode Guarantee Who picks the tool Use it when
"auto" May call a tool, may return plain text Model Never, when structured output is mandatory
"any" Must call a tool Model Several schemas, document type unknown — the tool name doubles as the classification
{"type": "tool", "name": "..."} Must call that tool You One particular extraction has to run, e.g. metadata before enrichment consumes it

The API has a fourth value, none, which blocks tool calls for that request.19 The guide does not test this — if it appears as an option, it is not the credited answer.

Strict schemas stop syntax errors, not semantic ones

This is the highest-value nuance in the statement. JSON Schema's strict mode is the feature name to know here. The guide lists it as strict mode for syntax error elimination — syntax errors, and nothing more. Note where it actually lives: strict is a field on the tool definition, not on the schema and not on tool_choice.48 Without it, the schema guides the model rather than binding it. A strict schema guarantees the shape: required keys present, values of the right type, enum values legal. It guarantees nothing about meaning. Line items can still fail to sum to the stated total; a vendor tax ID can land in the invoice-number field; a date can be wrong in perfect ISO 8601. Schema compliance is not correctness, and the remedy is semantic validation (4.4), not a stricter schema.

Schema design that does not induce fabrication

Two design decisions carry most of the weight:

  • Required vs optional. If a source document may legitimately not contain a field, model it as optional or nullable. A field marked required forces the model to produce something, and what it produces is a fabricated value. Making absence representable is what converts a hallucination into an honest null.49
  • Extensible enums. Closed enums break on real-world data. Add an "unclear" value for genuinely ambiguous cases, and an "other" value paired with a detail string so an unanticipated category can be captured rather than forced into the nearest wrong bucket.

Finally, a schema constrains structure but says nothing about formatting. Source documents write dates, currencies and units inconsistently. Put format normalization rules in the prompt alongside the strict schema: "dates as YYYY-MM-DD; amounts as decimals with no currency symbol; for a range, record the lower bound and note it".

tool_choice modes by guarantee, schema selection and fitGive the three tool_choice modes as rows against the three columns the exam tests — what the call guarantees, who selects the schema, and which stem it fits — so a stem maps to a mode in one lookup.guaranteeschema selectionfittool_choice autotool_choice autotool_choice anytool_choice anytool_choice forced tooltool_choice forced toolPlain text answer allowedPlain text answer allowedModel may skip every schemaModel may skip everyschemaUnsafe for mandatory structureUnsafe for mandatorystructureSome tool call guaranteedSome tool call guaranteedModel picks the schemaModel picks the schemaUnknown document typeUnknown document typeNamed tool call guaranteedNamed tool call guaranteedYou pick the schemaYou pick the schemaOne extraction must run firstOne extraction must run first
tool_choice modes by guarantee, schema selection and fit

Give the three tool_choice modes as rows against the three columns the exam tests — what the call guarantees, who selects the schema, and which stem it fits — so a stem maps to a mode in one lookup.

Extraction via tool use, and where each error class is caughtShow the concrete call path from document to validated record, and make clear that a strict tool schema removes syntax errors while only the semantic validator can catch wrong values.Extraction serviceExtractionserviceMessages APIMessagesAPISemantic validatorSemanticvalidatorDownstream systemDownstreamsystemdocument + extraction toolschema + tool_choicetool_use block,schema-compliant inputparsed record, syntaxhandled by strict modesemantic errors(sums, wrong field)validatedrecord
Extraction via tool use, and where each error class is caught

Show the concrete call path from document to validated record, and make clear that a strict tool schema removes syntax errors while only the semantic validator can catch wrong values.

An extraction tool schema designed against fabrication

Scenario 6 · Structured Data Extraction

Note what is not required: any field the document may legitimately omit is nullable, so "absent" has a legal representation. The currency enum ships with other plus a detail string and an unclear value, so an unanticipated or genuinely ambiguous currency does not get forced into USD.

json
{
  "name": "extract_invoice",
  "description": "Record the fields present in this invoice. Omit or null any field the document does not contain; never infer a value.",
  "input_schema": {
    "type": "object",
    "properties": {
      "invoice_number": { "type": "string" },
      "issue_date": {
        "type": ["string", "null"],
        "description": "YYYY-MM-DD, or null if the document has no issue date"
      },
      "vendor_tax_id": { "type": ["string", "null"] },
      "currency": {
        "type": "string",
        "enum": ["USD", "EUR", "GBP", "other", "unclear"]
      },
      "currency_detail": {
        "type": ["string", "null"],
        "description": "Required when currency is 'other'; the code as printed"
      },
      "line_items": {
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "description": { "type": "string" },
            "amount": { "type": "number" }
          },
          "required": ["description", "amount"]
        }
      }
    },
    "required": ["invoice_number", "currency", "line_items"]
  }
}
Tool definition used purely to constrain output shape — exam-shaped; a live `strict` tool also needs `strict: true` and `additionalProperties: false`

Choosing tool_choice: unknown document type vs mandatory step

Scenario 6 · Structured Data Extraction

Two different problems, two different modes. When the incoming document could be an invoice, a purchase order or a receipt, "any" guarantees structured output and tells you which schema matched via the returned tool name. When a downstream enrichment step requires metadata to already exist, force that specific tool by name. Leaving either case on "auto" risks a plain-text answer that breaks the pipeline.

typescript
// Document type unknown: the model must call one of the schemas,
// and its choice of tool name is the classification.
const res = await client.messages.create({
  model: 'claude-opus-5',
  max_tokens: 4096,
  tools: [extractInvoice, extractPurchaseOrder, extractReceipt],
  tool_choice: { type: 'any' },
  messages: [{ role: 'user', content: [documentBlock] }],
});

const call = res.content.find((b) => b.type === 'tool_use');
// call.name  -> which document type matched
// call.input -> the structured extraction, already schema-compliant

// Metadata must exist before the enrichment pass runs: force that one tool.
const meta = await client.messages.create({
  model: 'claude-opus-5',
  max_tokens: 2048,
  tools: [extractMetadata, extractInvoice],
  tool_choice: { type: 'tool', name: 'extract_metadata' },
  messages: [{ role: 'user', content: [documentBlock] }],
});
tool_choice selected per situation

Normalization rules in the prompt beside the strict schema

Scenario 6 · Structured Data Extraction

The schema says issue_date is a string and amount is a number. It cannot say that "11/02/26" is ambiguous, that "1.240,00" is European decimal notation, or that "about a tablespoon" is not a number. Those rules belong in the prompt, and they are what stops the model from silently guessing a convention.

text
Normalization rules:
- Dates: emit YYYY-MM-DD. If the source order is ambiguous (11/02/26) and the
  document gives no locale signal, set the date to null and set
  date_ambiguous to true. Do not pick an interpretation.
- Amounts: emit a decimal number with no currency symbol and no thousands
  separator. Treat "1.240,00" as 1240.00 only when the document uses comma
  decimals elsewhere; otherwise set the amount to null.
- Informal quantities ("about a tablespoon", "roughly room temperature"):
  leave the numeric field null and copy the phrase verbatim into the
  matching *_note field.
- Never convert, round, or infer a unit that the document does not state.
Prompt fragment paired with the extraction tool
  • Prose JSON requests: instructing the model to "respond with only valid JSON matching this shape" instead of defining a tool with a JSON schema because prose leaves syntax errors, markdown fences and preamble on the table while tool use eliminates them.
  • Leaving tool_choice on "auto" when structured output is mandatory instead of using "any" or a forced named tool because with "auto" the model may legitimately answer in plain text and the pipeline breaks.
  • All-required schemas: marking every field required so downstream code never sees a null, instead of making genuinely optional fields nullable because the model then fabricates a value to satisfy the schema.
  • Treating schema compliance as correctness: skipping semantic checks because the tool-use output validated, instead of verifying sums and field placement because a strict schema eliminates syntax errors and nothing else.
  • Map the stem to the tool_choice mode: "document type is unknown, several extraction schemas exist" → "any"; "this extraction must run before the enrichment step" → forced {"type":"tool","name":...}; "structured output is required" → anything but "auto".
  • Distractors here are prompt-side or post-processing fixes: "add 'return only JSON' to the prompt", "add a JSON repair/retry parser", "lower output length". The credited answer is tool use with a JSON schema.
  • If the stem describes valid JSON whose values are wrong (line items that do not sum, a value in the wrong field), the fault is semantic and the answer lives in 4.4 — not in a stricter schema or more required fields.
Beyond the exam — what the API does that this does not grade

strict: true is a top-level field on the tool definition, beside name and input_schema. It is not a schema keyword and not a tool_choice option. The schema must declare additionalProperties: false and required.48

Strict mode also drops several JSON Schema keywords: recursive $ref, numeric constraints (minimum, maximum, multipleOf) and string constraints (minLength, maxLength). Validate those in your own code. Schemas written for plain tool use routinely carry them — the CI schema in 3.6 uses "minimum": 1 — and under strict mode they stop being enforced.

There is also a second mechanism the guide does not cover. output_config: {format: {type: "json_schema", schema: ...}} constrains the response rather than a tool input, for the case where no tool call is wanted. It is incompatible with citations.47

References — 4 sources
  1. Define tools Anthropic The normative tool-definition reference — name, `description`, `input_schema` — with Anthropic's own guidance on description quality, and the place to check whether the unit's four description components are the source's four. It is also where `tool_choice` is enumerated: "there are four possible options", `auto`, `any`, `tool` and `none`. With `any` or `tool` the API prefills the assistant message, so no natural-language text precedes the `tool_use` block.
  2. Structured outputs Anthropic Documents a first-class structured-output path — `output_config.format`, formerly the `output_format` beta — that returns validated JSON directly without the extraction-tool trick, and explains how it differs from strict tool use. It also carries the supported Pydantic integration (`client.messages.parse()` with `output_format=Model`), which shows where schema validation stops and 4.4's semantic validator must begin.
  3. Strict tool use Anthropic The actual feature: `strict: true` on a tool definition, enforced by grammar-constrained sampling. It requires `additionalProperties: false` and `required`, and it drops several JSON Schema keywords — recursive `$ref`, `minimum`, `maximum`, `multipleOf`, `minLength`, `maxLength`.
  4. Reduce hallucinations Anthropic Anthropic's guidance on giving the model a legal way to say "not present", which is the general principle the nullable-field rule implements at the schema layer.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • Tool use (tool_use) with JSON schemas as the most reliable approach for guaranteed schema- compliant structured output, eliminating JSON syntax errors
  • The distinction between tool_choice: "auto" (model may return text instead of calling a tool), "any" (model must call a tool but can choose which), and forced tool selection (model must call a specific named tool)
  • That strict JSON schemas via tool use eliminate syntax errors but do not prevent semantic errors (e.g., line items that don't sum to total, values in wrong fields)
  • Schema design considerations: required vs optional fields, enum fields with "other" + detail string patterns for extensible categories

Skills in

  • Defining extraction tools with JSON schemas as input parameters and extracting structured data from the tool_use response
  • Setting tool_choice: "any" to guarantee structured output when multiple extraction schemas exist and the document type is unknown
  • Forcing a specific tool with tool_choice: {"type": "tool", "name": "extract_metadata"} to ensure a particular extraction runs before enrichment steps
  • Designing schema fields as optional (nullable) when source documents may not contain the information, preventing the model from fabricating values to satisfy required fields
  • Adding enum values like "unclear" for ambiguous cases and "other" + detail fields for extensible categorization
  • Including format normalization rules in prompts alongside strict output schemas to handle inconsistent source formatting
Back to top