Note on selection
Design human review workflows and confidence calibration
What you need to know
- A 97% aggregate can hide a document type or a field performing far worse; validate accuracy by document type and by field before reducing review.
- Confidence should be field-level, and thresholds must be calibrated on labeled validation sets rather than chosen intuitively.
- Route to humans on low confidence AND on ambiguous or contradictory source documents — high confidence over a self-contradicting document is still a review case.
- Stratified random sampling of high-confidence extractions surfaces novel error patterns; reweight per-stratum rates before quoting an overall error rate.
- The goal is prioritizing limited reviewer capacity toward the highest expected error, not reviewing more overall.
Human review is a scarce resource, so the design question is how to spend it where the errors actually are — and how to know where that is.
Aggregate accuracy hides the failures that matter. A pipeline reporting 97% overall accuracy tells you almost nothing about risk, because the mean is dominated by the high-volume, easy segment. Underneath it a specific document type (handwritten forms, low-quality scans, a foreign-language variant) or a specific field (tax ID, line-item totals, dates in ambiguous formats) may be running far below the headline. Before reducing human review, analyze accuracy by document type and by field and confirm performance is consistent across all segments. Automating on the aggregate is the classic wrong answer: it automates precisely the segment where automation fails. This is measured, not hypothetical. Commercial gender classifiers with acceptable headline accuracy ran 0.8% error on lighter-skinned males and 34.7% on darker-skinned females.67 Disaggregated reporting is a norm rather than an optional analysis, and a federal risk framework states it plainly: accuracy measurements may include disaggregation of results for different data segments.6869
Field-level confidence, calibrated against labeled data. Have the model output confidence per extracted field rather than one score per document. A document can have nine solid fields and one guess, so the useful routing decision is field-scoped. A verbalized score can be made to track probabilities, which is why eliciting one is worth doing at all.61 But raw scores are not trustworthy on their own. They must be calibrated using labeled validation sets until a threshold means something empirically. A calibration run might conclude, for instance, that "extractions above 0.92 have a measured 0.4% error rate on this document type". The threshold is an output of measurement, not a number someone chose. The named post-hoc method is temperature scaling, fitted against held-out labeled data.72 Building that labeled set is its own task, and it must cover the ambiguous cases where even humans struggle to reach consensus.65
Route on two axes, not one. Confidence decides nothing by itself, because the score describes the model and the second axis describes the document:
| Confidence against the calibrated threshold | Source evidence | Route | Why |
|---|---|---|---|
| Above | Clear and self-consistent | Auto-approve, then sample it | The measured error rate is acceptable; sampling is what keeps that true |
| Above | Ambiguous or contradictory | Human review | The model is confident about a value the document does not settle |
| Below | Clear and self-consistent | Human review | Below the calibrated threshold means an error rate nobody accepted |
| Below | Ambiguous or contradictory | Priority human review | Highest expected error — spend the scarcest capacity here first |
Row 2 is the one exam items are built on. A high score means the model read the document consistently, not that the document says one thing: a page stating two different totals reaches a person whatever the score says. Prioritizing along both axes is what lets limited reviewer capacity cover the highest expected error. The deferral literature supports the general move: a pure confidence cutoff is not the right hand-off rule.73 Read it honestly, though — its second axis is the human reviewer's competence, not document ambiguity. Ambiguity becomes checkable when each extracted value is bound to a verifiable span of the source document.66
Keep measuring what you automated. High-confidence extractions that skip review still need stratified random sampling: sample across strata (document type, field, confidence band) rather than uniformly, so rare-but-risky segments are actually represented. Stratification is what lets you control precision within a stratum rather than only overall.7071 This gives an ongoing error-rate estimate for the automated path — provided the per-stratum rates are recombined weighted by each stratum's share of that path. The sampler deliberately over-draws from small strata, so the raw pooled error rate across the sample is not the population error rate. It over-weights the rare segments it exists to surface. Detecting a novel pattern and estimating the overall rate want opposite allocations. The weighting is what lets one draw serve both. Sampling also detects novel error patterns that appear when document formats change upstream. Without sampling, a new failure mode is invisible until it reaches a customer, because everything the model was confident about was accepted unexamined.
Show that routing is two-dimensional: confidence alone does not decide, because source ambiguity overrides a high score, and the automated cell still owes a sampled audit.
Click a cell to reveal the reason its route is what it is.
Make the masking effect concrete so the learner instinctively decomposes any headline accuracy figure by document type and field.
Worked examples
Segment analysis before switching off review
Scenario 6 · Structured Data ExtractionThe team wants to auto-approve high-confidence extractions. The aggregate says 97.1%. Breaking accuracy down by document type and field shows two segments that must keep full review, and the decision becomes segment-scoped rather than global.
document type volume accuracy decision
--------------------------------------------------------------
typed PDF invoice 62% 99.2% eligible for automation
digital receipt 21% 98.4% eligible for automation
scanned invoice 12% 96.1% automate, sample at 5%
handwritten form 5% 71.3% keep 100% human review
field accuracy decision
--------------------------------------------------------------
invoice_number 99.4% automate
total_amount 98.9% automate
tax_id 82.0% keep human review (all doc types)
line_item_dates 91.2% automate only for typed PDFs
Aggregate: 97.1% <- would have justified automating everything.Field-level confidence with a calibrated threshold
Scenario 6 · Structured Data ExtractionThe extractor emits a confidence value per field plus flags describing the source evidence. Routing uses the calibrated threshold for that field and document type, and any source ambiguity overrides a high score.
{
"document_id": "inv-2026-08-0417",
"document_type": "scanned_invoice",
"fields": {
"invoice_number": { "value": "INV-88213", "confidence": 0.97,
"source_span": "p1:header", "source_flags": [] },
"total_amount": { "value": "1284.50", "confidence": 0.94,
"source_span": "p2:total-row",
"source_flags": ["two_conflicting_totals"] },
"tax_id": { "value": "DE811234567", "confidence": 0.61,
"source_span": "p1:footer", "source_flags": ["low_ocr_quality"] }
},
"routing": {
"invoice_number": "auto_approve",
"total_amount": "human_review (source conflict overrides confidence)",
"tax_id": "human_review (below calibrated threshold 0.92)"
}
}Stratified sampling of the automated path
Scenario 6 · Structured Data ExtractionAuto-approved extractions are not left unexamined. A sampler draws from each stratum — document type crossed with field and confidence band — so low-volume, high-risk strata are represented instead of being swamped by the dominant one. The sampled records go to reviewers as labeled data, feeding both the reweighted error-rate estimate and threshold recalibration.
type Stratum = { docType: string; field: string; band: '0.92-0.95' | '0.95-1.0' };
// Uniform sampling would spend ~62% of the budget on typed PDFs.
// Stratified sampling guarantees coverage of every risky stratum.
function stratifiedSample(pool: Extraction[], budget: number) {
const strata = groupBy(pool, keyOf); // docType x field x band
const perStratum = Math.max(5, Math.floor(budget / strata.size));
return [...strata.values()].flatMap((rows) =>
randomDraw(rows, Math.min(perStratum, rows.length)),
);
}
// Reviewer verdicts feed two loops:
// 1. error rate per stratum, reweighted by stratum share of the automated
// path -> is that path still within tolerance overall?
// 2. new failure descriptions -> novel error patterns (e.g. a changed
// vendor template) that no existing validation rule would have caught.
// Do NOT pool the raw sample: equal-per-stratum allocation over-weights
// the rare strata by design.Anti-patterns
- Reducing human review on the strength of an aggregate accuracy figure instead of per-document-type and per-field analysis because the mean is dominated by the easy segment and hides the one that fails.
- Auto-approving everything above a confidence threshold without stratified sampling of that population because novel error patterns from changed document formats then stay invisible until a customer finds them.
- Asking for human review with no confidence signal or priority ordering because undifferentiated queues spend scarce reviewer capacity on the extractions least likely to be wrong.
- Trusting raw model confidence as a probability instead of calibrating thresholds against a labeled validation set because an uncalibrated score does not map to any measured error rate.
How it is examined
- Any stem quoting a single impressive accuracy number (97%, 99%) is testing aggregate masking — the credited answer breaks accuracy down by document type and field before automating.
- Watch the difference between "sample randomly" and "sample stratified". Uniform sampling is the plausible distractor; stratification is what surfaces rare segments and novel patterns.
- Routing items usually have two triggers hidden in the stem: low confidence and ambiguous or contradictory source documents. An option that only handles low confidence is incomplete.
References — 10 sources
- Teaching Models to Express Their Uncertainty in Words Lin, Hilton & Evans, TMLR 2022 Verbalized confidence maps to well-calibrated probabilities and stays moderately calibrated under distribution shift, grounded in latent representations that correlate with epistemic uncertainty.
- Define success criteria and build evaluations Anthropic How to build the labeled evaluation set calibration presupposes, including deliberate edge-case coverage of "ambiguous test cases where even humans would find it hard to reach an assessment consensus".
- Citations Anthropic Model-generated citations bound to spans of a supplied document, returned as structured blocks — the artifact that makes an ambiguity flag checkable rather than a judgment call.
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification Buolamwini & Gebru, PMLR 81 (FAT* 2018) The canonical measured instance of aggregate masking: 0.8% error for lighter-skinned males against 34.7% for darker-skinned females, in systems whose headline accuracy looked acceptable.
- Model Cards for Model Reporting Mitchell et al., FAT* 2019 Establishes disaggregated evaluation as a reporting norm rather than an optional analysis: benchmarked performance broken out by group and by intersection of groups.
- NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0) NIST A federal framework stating the segment rule directly: "Accuracy measurements may include disaggregation of results for different data segments", paired with representative test sets and human intervention where the system cannot correct itself.
- Survey Methods and Practices (Catalogue no. 12-587-X) Statistics Canada The national statistical agency's survey-design handbook — the standard reference for stratified sampling, covering sample size, allocation across strata, and selection.
- Section 2. Stratified sampling (Survey Methodology 45(2), Cat. 12-001-X) Statistics Canada The load-bearing property, stated exactly: stratification "enables controlling sample sizes and precision of estimates for the strata", and can improve overall precision for a fixed cost.
- On Calibration of Modern Neural Networks Guo, Pleiss, Sun & Weinberger, ICML 2017 Raw model confidence is not trustworthy as a probability, and temperature scaling is the post-hoc procedure that fixes it against held-out labeled data. Measured on image and document classifiers, not on verbalized LLM self-report.
- Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer Madras, Pitassi & Zemel, NeurIPS 2018 The literature on when a model should hand off to a human. Supports "one axis is insufficient"; its second axis is the human reviewer's competence, not document ambiguity.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- The risk that aggregate accuracy metrics (e.g., 97% overall) may mask poor performance on specific document types or fields
- Stratified random sampling for measuring error rates in high-confidence extractions and detecting novel error patterns
- Field-level confidence scores calibrated using labeled validation sets for routing review attention
- The importance of validating accuracy by document type and field segment before automating high- confidence extractions
Skills in
- Implementing stratified random sampling of high-confidence extractions for ongoing error rate measurement and novel pattern detection
- Analyzing accuracy by document type and field to verify consistent performance across all segments before reducing human review
- Having models output field-level confidence scores, then calibrating review thresholds using labeled validation sets
- Routing extractions with low model confidence or ambiguous/contradictory source documents to human review, prioritizing limited reviewer capacity