Note on selection
Implement validation, retry, and feedback loops for extraction quality
What you need to know
- A retry without the specific validation error in the prompt is just a re-roll; append the errors so the model can self-correct.
- The follow-up request carries the original document, the failed extraction, and the specific validation errors.
- Retries fix format and structural errors; they cannot supply information that is absent from the source document — escalate instead.
- A strict tool schema removes syntax errors, so your validator exists specifically for semantic errors: values that do not sum, values in the wrong field.
- Pydantic is the guide-named library for this loop: it handles schema validation while your own checks raise semantic errors — schema-valid is not the same as semantically valid, and either failure feeds the same validation-retry loop.
- Have the extraction emit calculated_total next to stated_total, and a conflict_detected boolean, so inconsistency is a comparison rather than a guess.
- Tag findings with detected_pattern so dismissals can be analyzed by the code construct that triggered them.
Retry only works if the model learns what failed
A retry that re-sends the original prompt unchanged is a re-roll: same inputs, same distribution, no reason to expect a different outcome. Retry-with-error-feedback means the follow-up request carries three things:21
- the original document,
- the failed extraction exactly as the model produced it, and
- the specific validation errors — not "invalid output" but "line_items sum to 1240.00 but
stated_totalis 1420.00" and "vendor_tax_id contains what appears to be an invoice number".
With those three, the second attempt is a self-correction task on a concrete defect rather than a fresh guess.3
Know which failures a retry can fix
Classifying the failure comes before choosing the response, and the guide treats that classification as its own skill.
| Failure kind | What it looks like | Retry with error feedback? | What to do |
|---|---|---|---|
| Syntax | a markdown fence, a trailing comma, an unquoted key | not applicable | already eliminated upstream by tool use with a strict JSON schema (4.3) |
| Format or structural | a date in the wrong notation, a nested object flattened, a normalization rule ignored | yes — this is the class retry exists for | resend the three items above, on a bounded loop |
| Semantic | line items that do not sum to the stated total, a vendor tax ID sitting in the invoice-number field | yes | same loop, driven by your validator's errors |
| Information absent from the source | a figure that lives only in a referenced external appendix you never supplied | no | escalate, or fetch the missing source and re-extract |
No amount of re-asking creates data the model never saw. Retrying an absent-information failure just multiplies cost and, worse, pressures the model toward fabrication.
Semantic vs syntax errors
The semantic row is the one your validator exists for: values that do not reconcile, values in the wrong field, internally contradictory records. Because the schema cannot catch these, your validator is the only thing standing between a plausible-looking record and a corrupted downstream system.
Pydantic is the guide's named example of the validator in this loop. You declare the record as a Pydantic model and it does the schema validation, while your own field and cross-field checks raise the semantic validation errors. The guide's phrasing is "when Pydantic or JSON schema validation fails, send a follow-up request".47 The same validate → error → retry loop is what you build regardless of which of the two raised the failure. Keep the two levels distinct: schema-valid means the shape parses and the types match; semantically valid means the values are true of the source document.
Design the schema to help you validate
The strongest pattern is to make the model surface the evidence for its own consistency check:
- Extract
calculated_total(sum the line items yourself, in the extraction) alongsidestated_total(what the document prints). A mismatch is then a single comparison rather than a re-derivation. - Add a
conflict_detectedboolean (with a note) so the model can report that the source is internally inconsistent, instead of silently picking one of two contradictory values.
Feedback loops beyond a single record
The same principle applies to review findings. Adding a detected_pattern field — the code construct that triggered the finding — turns individual dismissals into analyzable data. When developers dismiss findings, you can group by detected_pattern and see that, say, 90% of dismissals come from one construct, which points straight at the criteria to rewrite (4.1) or the negative few-shot example to add (4.2). Without that field you have a dismissal rate and no idea what to fix.
Show that the retry edge carries the document, the failed extraction and the specific errors, that the loop is bounded, and that an absent-information failure exits to escalation rather than looping.
Click a transition to highlight its label — the exit condition it tests, and for the retry edge the document, failed extraction and specific errors it carries.
Worked examples
A retry loop that feeds the validation errors back
Scenario 6 · Structured Data ExtractionThe loop is bounded, the errors are part of the next prompt, and the terminal branch distinguishes "the model got the shape wrong" from "the data is not in this document". Only the first is worth retrying; the second is escalated with the errors attached so a human sees why.
let extraction = await extract(document); // tool_use, schema-compliant
let errors = validateSemantics(extraction); // sums, field placement, conflicts
for (let attempt = 1; attempt <= 2 && errors.length > 0; attempt++) {
// The follow-up request includes the document, the failed extraction,
// and the specific errors -- this is what makes it a correction, not a re-roll.
extraction = await extract(document, {
failedExtraction: extraction,
validationErrors: errors,
// e.g. "line_items sum to 1240.00 but stated_total is 1420.00"
// "vendor_tax_id looks like an invoice number, not a tax id"
});
errors = validateSemantics(extraction);
}
if (errors.length > 0) {
// Retrying cannot invent data the document never contained.
return errors.every(isInformationAbsentFromSource)
? escalateToHuman(document, extraction, errors)
: quarantine(document, extraction, errors);
}
return extraction;Self-correction fields in the extraction schema
Scenario 6 · Structured Data ExtractionRather than recomputing totals outside the model and hoping the field mapping was right, the schema asks for both numbers plus an explicit conflict flag. The validator's job collapses to comparing two fields and reading one boolean, and a genuinely inconsistent source document is reported as such instead of being quietly resolved.
{
"stated_total": {
"type": ["number", "null"],
"description": "The total as printed on the document, verbatim"
},
"calculated_total": {
"type": "number",
"description": "The sum of line_items[].amount that you extracted"
},
"conflict_detected": {
"type": "boolean",
"description": "true when the source document itself is internally inconsistent"
},
"conflict_note": {
"type": ["string", "null"],
"description": "Required when conflict_detected is true: which values disagree"
}
}detected_pattern turns dismissals into a feedback loop
Scenario 5 · Claude Code for Continuous IntegrationEvery finding records the construct that triggered it. After a week of PRs the dismissal data can be grouped by detected_pattern, which is how you discover that one construct accounts for most of the noise — the input to rewriting that category's criteria or adding a negative few-shot example.
{
"location": "src/billing/invoice.ts:88",
"category": "correctness",
"severity": "HIGH",
"issue": "total computed before discount is applied",
"suggested_fix": "compute subtotal, apply discount, then sum tax",
"detected_pattern": "reduce_over_items_without_discount_call",
"dismissed": true,
"dismissal_reason": "discount applied by caller in applyPromotion()"
}Anti-patterns
- Blind retry: re-sending the identical prompt after a validation failure instead of appending the specific validation errors because nothing in the second request tells the model what to correct.
- Retrying for absent data: burning retries when the required value exists only in an external document you never supplied, instead of escalating or fetching the source because retry cannot create information the model never saw — and pressure to answer invites fabrication.
- Treating tool-use schema compliance as validation: skipping semantic checks because the output parsed, instead of verifying sums and field placement because tool use eliminates syntax errors only.
- Free-text findings with no detected_pattern field, instead of tagging each finding with the construct that triggered it because you cannot analyze dismissal patterns you never recorded.
How it is examined
- When a stem says an extraction failed the same way after three retries and the missing figure lives in a document that was not provided, the credited answer is that retries are ineffective here — escalate or supply the source. "Increase the retry count" and "add few-shot examples" are the distractors.
- When the stem says the JSON was valid but the line items do not sum to the total, look for semantic validation plus retry-with-error-feedback. "Make the schema stricter" or "mark more fields required" are wrong — the schema already did its job.
- Options that merely log validation failures for later analysis are distractors against the option that includes the errors in the follow-up request. Logging is not feedback.
References — 3 sources
- Handle tool calls Anthropic The exact `tool_result` message contract, including the case where text placed before `tool_result` blocks ends the turn early and returns a 400.
- Troubleshooting tool use Anthropic The documented diagnostic order for wrong-tool-selected symptoms — where a reader checks whether auditing the system prompt for keyword-sensitive wording really is the recommended first move.
- Structured outputs Anthropic Documents a first-class structured-output path — `output_config.format`, formerly the `output_format` beta — that returns validated JSON directly without the extraction-tool trick, and explains how it differs from strict tool use. It also carries the supported Pydantic integration (`client.messages.parse()` with `output_format=Model`), which shows where schema validation stops and 4.4's semantic validator must begin.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Retry-with-error-feedback: appending specific validation errors to the prompt on retry to guide the model toward correction
- The limits of retry: retries are ineffective when the required information is simply absent from the source document (vs format or structural errors)
- Feedback loop design: tracking which code constructs trigger findings (detected_pattern field) to enable systematic analysis of dismissal patterns
- The difference between semantic validation errors (values don't sum, wrong field placement) and schema syntax errors (eliminated by tool use)
Skills in
- Implementing follow-up requests that include the original document, the failed extraction, and specific validation errors for model self-correction
- Identifying when retries will be ineffective (e.g., information exists only in an external document not provided) versus when they will succeed (format mismatches, structural output errors)
- Adding detected_pattern fields to structured findings to enable analysis of false positive patterns when developers dismiss findings
- Designing self-correction validation flows: extracting "calculated_total" alongside "stated_total" to flag discrepancies, adding "conflict_detected" booleans for inconsistent source data