Skip to content
CCAR-FAcademy
Domain 4 · Statement 4.2 2 of 6
4.2

Apply few-shot prompting to improve output consistency and quality

  • When detailed instructions still produce inconsistent output, few-shot examples are the highest-leverage next move — not longer instructions.
  • Use 2–4 targeted examples chosen at the ambiguous boundary; a pile of easy cases teaches the model nothing about where the line is. 2–4 is the few-shot count; the 2–3 count belongs to iterative-refinement input/output examples (3.5).
  • An ambiguous-case example teaches nothing without the reasoning that rejected the plausible alternative.
  • Demonstrate the exact output fields (location, issue, severity, suggested fix) instead of describing them in prose.
  • Pair a genuine issue with a superficially similar acceptable pattern to cut false positives while preserving generalization.
  • In extraction, one example per document structure variant is the stated fix for hallucination and for empty/null required fields.

Few-shot is the fix when instructions alone are inconsistent

The guide is unusually direct here: few-shot examples are described as the most effective technique for getting consistently formatted, actionable output when detailed instructions alone produce inconsistent results. The mechanism is that an example demonstrates the target simultaneously along every axis you care about — field set, field order, granularity, tone, length, and what a "good" value looks like — whereas prose has to describe each axis separately and each description is itself open to interpretation.46 If your first instinct on inconsistent output is "write longer instructions", the exam expects the second instinct: show two to four examples.

Examples earn their place at the ambiguous boundary

The target count is small — 2 to 4 targeted examples — which means selection matters more than volume.41 Keep this count attached to its scope: 2–4 is the few-shot prompting number, for examples that teach ambiguous-case judgment inside a prompt. The sibling number you will also see on the exam is 2–3, which belongs to iterative refinement (3.5): concrete input/output pairs supplied to pin down a transformation the model keeps interpreting differently. Same idea, different statement and different job, so read which one the stem is describing before you pick a count. The examples that pay for themselves are the ambiguous ones: which tool to select for a request that plausibly maps to two tools; whether a partially covered branch counts as a test coverage gap; whether an unusual-looking code pattern is a genuine defect or an accepted project convention. Crucially, an ambiguous-case example should include the reasoning for why one action was chosen over a plausible alternative. That reasoning is the part that transfers: it teaches the model the principle behind the boundary rather than a single mapping, which is why well-chosen few-shot examples let the model generalize judgment to novel patterns instead of only matching the cases you pre-specified.

Format pinning and false positive reduction

Two concrete jobs show up repeatedly:

  • Format consistency. Demonstrate the exact output shape you want — location, issue, severity, suggested fix — in the example, rather than describing those fields in prose.
  • False positive reduction with generalization. Include an example of an acceptable pattern that must not be flagged alongside a superficially similar genuine issue. This teaches the exclusion boundary. Without a negative example the model can only guess where the line sits, and the guide notes that this pairing reduces false positives while still allowing generalization — you are not enumerating an allowlist.

Extraction: examples reduce hallucination

In extraction tasks, few-shot examples are the stated remedy for hallucination and for empty or null values in required fields. Documents differ structurally: citations may be inline or collected in a bibliography; methodology may have its own section or be buried in a results paragraph; measurements may be informal ("about a tablespoon"). An example per structural variant shows the model where to look and how to normalize, instead of letting it either invent a plausible value or give up and return nothing. Pair this with a nullable schema (4.3) so "genuinely absent" has a legal representation.

Few-shot for an ambiguous tool-selection decision

Scenario 1 · Customer Support Resolution Agent

A support request that mentions both a charge and a delivery problem plausibly maps to lookup_order or process_refund. Instructions alone ("pick the most appropriate tool") leave it a coin flip. The example below shows the choice and the reasoning that rejected the alternative, which is what lets the model handle a novel ambiguous request the same way.

text
User: "I was charged twice for order 4471 and it still hasn't arrived."

Reasoning: Two issues are present. The duplicate charge is asserted, not
verified — process_refund would act on an unverified claim and is not
reversible. lookup_order returns both the shipment state and the charge
history, so it resolves which issue is real before any irreversible action.
Choose the read before the write.

Action: lookup_order(order_id="4471")

---

User: "Cancel my subscription and refund last month."

Reasoning: Two requests, both explicit and both write actions. Refund policy
depends on subscription state at time of cancellation, so cancellation must
be sequenced first. No read is needed to disambiguate intent.

Action: cancel_subscription(...) then process_refund(...)
One of 2–4 ambiguous-case examples

Format-pinning plus an acceptable-pattern negative example

Scenario 5 · Claude Code for Continuous Integration

The review bot's output shape was drifting: sometimes a paragraph, sometimes a bullet list, sometimes without a line number, which broke the code that turned findings into PR comments. Two examples fix the shape, and the second doubles as the exclusion boundary — it shows a pattern that looks like the flagged one but must be skipped, so precision improves without hard-coding an allowlist.

markdown
EXAMPLE 1 — report
  location: src/billing/invoice.ts:88
  issue: total is computed from items before the discount is applied, so
    discounted invoices overcharge by the discount amount
  severity: HIGH
  suggested_fix: compute subtotal, apply discount, then sum tax

EXAMPLE 2 — skip (looks similar, is not a defect)
  code: const total = items.reduce(sum, 0)   // discount applied upstream
  finding: none
  why: the discount is applied by the caller in applyPromotion(); flagging
    this would be a false positive. Only flag when no caller in the diff
    applies the discount.
Format demonstration + negative example

Examples covering varied document structures

Scenario 6 · Structured Data Extraction

The extraction pipeline was returning null for citations on roughly a third of papers, and occasionally inventing a reference list. The papers were not malformed — they simply put citations in different places. Adding one example per structural variant addressed both the nulls and the fabrication, because each example demonstrates where the information lives in that layout.

text
VARIANT A — inline citations, no bibliography
  Source: "...consistent with earlier work (Okonkwo & Reyes, 2019)."
  Extract: citations = [{ authors: "Okonkwo & Reyes", year: 2019,
                          title: null }]
  Note: title is absent from the document -> null, do NOT infer one.

VARIANT B — numeric markers plus a References section
  Source: "...as shown in [3]."  /  "[3] Okonkwo A. Signal drift. 2019."
  Extract: citations = [{ authors: "Okonkwo A.",
                          title: "Signal drift", year: 2019 }]
  Note: resolve the marker against the References section before extracting.

VARIANT C — methodology embedded in the results narrative
  Source: "Samples were run in triplicate at roughly room temperature..."
  Extract: method.replicates = 3,
           method.temperature_c = null,
           method.temperature_note = "roughly room temperature"
  Note: informal measurements go in the note field; never convert them to
        a number.
Structure-variant examples for a citation extractor
  • Easy-case few-shot: filling the example block with unambiguous cases instead of the boundary cases the model actually gets wrong because examples teach where the line is and easy examples put the line nowhere useful.
  • Answer-only examples: showing the chosen action without the reasoning that rejected the plausible alternative because the model then copies the surface of the example instead of generalizing the judgment to novel patterns.
  • Describing the format instead of demonstrating it: specifying "include location, issue, severity and fix" in prose rather than showing a formatted example because prose specifications drift across runs while a demonstrated shape does not.
  • Positive-only examples: never showing an acceptable pattern that must not be flagged because the model cannot infer the exclusion boundary and keeps producing false positives on look-alike code.
  • A stem that says "detailed instructions produce inconsistent output" is signaling few-shot. Distractors will offer more explicit instructions, a longer rules list, or a stricter tone — the credited answer adds 2–4 examples.
  • Watch the example count and the example choice. "Add 20 examples covering every case" and "add one example of the typical case" are both wrong shapes; the guide asks for 2–4 targeted examples of the ambiguous cases.
  • If a stem describes empty or null values in required extraction fields across differently formatted documents, the credited fix is adding examples of correct extraction from those formats — often combined with making the fields nullable (task statement 4.3).
References — 2 sources
  1. Prompting best practices Anthropic The consolidated reference on examples, clarity and structure. It recommends 3–5 examples for best results, which is neither of the counts this corpus separates — confirming that 2–3 here and 2–4 in 4.2 are guide-local figures to be reproduced on the exam, not Anthropic guidance.
  2. Increase output consistency Anthropic The documented levers that do move consistency — exact output formats, response prefill, constraining with examples, retrieval grounding — which is the positive list behind 4.1's negative claim about "be conservative". It also states 4.2's central mechanism plainly: "Provide examples of your desired output. This is more effective than abstract instructions."
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • Few-shot examples as the most effective technique for achieving consistently formatted, actionable output when detailed instructions alone produce inconsistent results
  • The role of few-shot examples in demonstrating ambiguous-case handling (e.g., tool selection for ambiguous requests, branch-level test coverage gaps)
  • How few-shot examples enable the model to generalize judgment to novel patterns rather than matching only pre-specified cases
  • The effectiveness of few-shot examples for reducing hallucination in extraction tasks (e.g., handling informal measurements, varied document structures)

Skills in

  • Creating 2-4 targeted few-shot examples for ambiguous scenarios that show reasoning for why one action was chosen over plausible alternatives
  • Including few-shot examples that demonstrate specific desired output format (location, issue, severity, suggested fix) to achieve consistency
  • Providing few-shot examples distinguishing acceptable code patterns from genuine issues to reduce false positives while enabling generalization
  • Using few-shot examples to demonstrate correct handling of varied document structures (inline citations vs bibliographies, methodology sections vs embedded details)
  • Adding few-shot examples showing correct extraction from documents with varied formats to address empty/null extraction of required fields
Back to top