Skip to content
CCAR-FAcademy
Domain 4 · Statement 4.1 1 of 6
4.1

Design prompts with explicit criteria to improve precision and reduce false positives

  • Replace vague quality instructions with decidable categorical criteria the model can apply the same way twice.
  • "Be conservative" and "only report high-confidence findings" are not precision levers — they change tone, not the decision boundary.
  • Precision is per-category but trust is global: one noisy category undermines confidence in the accurate ones.
  • Temporarily disabling a high-false-positive category is a valid way to restore trust while you improve its prompt.
  • Anchor every severity level to a concrete code example so classification is repeatable.
  • State both axes separately: what counts as a finding, and how severe it is. Neither implies the other.

Explicit criteria are a decision rule; vague instructions are a wish

The exam's canonical contrast is "check that comments are accurate" versus "flag comments only when claimed behavior contradicts actual code behavior".41 The second one is not merely more words — it is a decidable predicate. The model can apply it to a comment/code pair and get the same answer twice. The first one forces the model to invent its own bar for "accurate", and it will invent a slightly different bar on every file, every run, and every PR. Inconsistent classification is the direct consequence of a criterion the model has to define for itself.

Why "be conservative" does not improve precision

Instructions like "be conservative" or "only report high-confidence findings" are meta-instructions about the model's internal certainty. They do not change what counts as a finding; they only ask the model to feel differently about findings it was already going to produce.46 There is no calibrated confidence threshold behind such a request, so the observable effect is usually tone and hedging rather than a moved decision boundary. The lever that actually moves precision is categorical: name the issue classes that are in scope (bugs, security vulnerabilities, correctness defects) and the classes that are explicitly out of scope (minor style, formatting, project-local patterns the team has deliberately adopted).

False positives are contagious

The reason precision matters more than raw recall in a CI review bot is a trust effect: a category with a high false positive rate undermines developer confidence in the categories that are accurate. Once a reviewer has dismissed eight of ten "possible null dereference" comments, they stop reading the security comments too. Precision is therefore a per-category property with a global blast radius.

That gives you two distinct engineering moves:

  • Temporarily disable the offending category while you improve its prompt. You trade coverage you were not getting value from anyway for restored trust in the remaining categories. This is a legitimate, exam-relevant answer — not a cop-out.
  • Define explicit severity criteria with a concrete code example for each level. Naming levels "critical / major / minor" is another vague criterion. Anchoring each level to a short code sample ("critical: an unchecked index into a caller-supplied array") makes classification consistent across runs and reviewers.45

Reporting criteria (report vs skip) and severity criteria (how bad) are separate axes. Both need to be explicit, and neither should be delegated to the model's sense of confidence.

Vague instruction vs explicit categorical criteriaShow that precision comes from replacing the decision rule with a decidable predicate, not from asking the model to be more confident or conservative — and show the downstream trust consequence of each path.Vague instructionVagueinstructionExplicit categorical criteriaExplicitcategoricalcriteriaModel invents its own barModel invents itsown barModel applies a fixed testModel applies afixed testInconsistent classificationInconsistentclassificationConsistent classificationConsistentclassificationHigh false positive rateHigh falsepositive rateFindings acted onFindings acted onTrust erodes across all categoriesTrust erodesacross allcategoriescheck comments are accurate,no decidable predicateflag only on codecontradiction, decidablepredicatevaries per runrepeatablenoisesignalspillover
Vague instruction vs explicit categorical criteria

Show that precision comes from replacing the decision rule with a decidable predicate, not from asking the model to be more confident or conservative — and show the downstream trust consequence of each path.

Rewriting a comment-accuracy check for a CI review bot

Scenario 5 · Claude Code for Continuous Integration

The first version of the prompt produced comments on every docstring that was merely terse or slightly out of date, and developers stopped reading the bot. The rewrite does not add "be careful" — it replaces the criterion with a contradiction test plus an explicit skip list.

markdown
BEFORE (vague — the model invents the bar)
  Check that comments are accurate.

AFTER (explicit categorical criterion)
  Flag a comment ONLY when the behavior it claims contradicts the behavior of
  the code it documents. A contradiction means: the comment states a return
  value, side effect, precondition, or error case that the code does not
  implement.

  Report:
    - comment says "returns null on failure", code throws
    - comment says "thread-safe", method mutates shared state without a lock
    - comment documents a parameter that no longer exists

  Skip:
    - comments that are terse, informal, or stylistically inconsistent
    - missing comments
    - comments that are incomplete but not contradictory
    - TODO/FIXME notes
Before / after criterion for the comment-accuracy category

Severity levels anchored to concrete code examples

Scenario 5 · Claude Code for Continuous Integration

Qualitative labels drift between runs. Giving one short code example per level turns severity into a matching task instead of a judgment call, which is what makes the classification consistent enough to drive routing (block the merge vs post a comment).

markdown
BLOCKING — a defect that can corrupt data or bypass authorization.
  Example:
    query = "SELECT * FROM users WHERE id = " + req.params.id

HIGH — a defect that produces incorrect behavior on a reachable input path.
  Example:
    for (let i = 0; i <= items.length; i++) total += items[i].price

MEDIUM — a defect that only manifests under a documented edge case.
  Example:
    parseInt(userInput)   // no radix, no NaN check

Anything that does not match one of the examples above is NOT reported.
Severity rubric fragment

Temporarily disabling a high-false-positive category

Scenario 5 · Claude Code for Continuous Integration

Telemetry shows developers dismiss 80% of findings in the race_condition category, and dismissal rates are rising in security too — the noise is bleeding across categories. The fix is to switch the noisy category off in the pipeline config while its criteria are rewritten, then re-enable it behind a measured dismissal rate. Turning the whole bot off, or leaving the category on and asking the model to be more careful, are both worse.

yaml
review_categories:
  correctness:      { enabled: true }
  security:         { enabled: true }
  race_condition:
    enabled: false
    # Disabled 2026-02-11: 80% dismissal rate was eroding trust in the
    # correctness and security categories. Re-enable when the rewritten
    # criteria hold dismissal under 20% on the replay corpus.
  style:            { enabled: false }   # explicitly out of scope
CI review configuration during remediation
  • Confidence-based filtering: instructing the model to "only report high-confidence findings" instead of naming which categories to report and which to skip because an instruction about certainty gives the model no decision rule and leaves the boundary exactly where it was. Note the scope of this ban: it forbids confidence as an upstream filter deciding what counts as a finding. Attaching self-reported confidence to an already-generated finding as downstream routing metadata is a different and credited technique (4.6).
  • Vague quality directives: asking the model to "check that comments are accurate" instead of defining the contradiction test because the model must then invent the bar and will invent a different one on each run.
  • Unanchored severity labels: naming levels critical/major/minor without a concrete code example per level because the model reclassifies the same defect differently across files and the routing built on top of it becomes unreliable.
  • Shipping a known-noisy category: leaving a high-false-positive category enabled while you iterate instead of disabling it temporarily because its noise erodes developer trust in the categories that are accurate.
  • Scenario 5 stems typically describe a review bot whose findings are being dismissed and ask what to change. Options that adjust how the model should feel about a finding (be conservative, be careful, only high confidence) are distractors; the credited option changes what counts as a finding.
  • When a stem says one category is noisy and overall trust is falling, the expected answer pairs a temporary disable of that category with prompt improvement — not accepting the noise, and not disabling review entirely.
  • If an option offers "add explicit severity definitions" versus "add explicit severity definitions with a code example per level", the version with concrete examples is the credited one — the guide names consistent classification as the goal.
References — 3 sources
  1. Prompting best practices Anthropic The consolidated reference on examples, clarity and structure. It recommends 3–5 examples for best results, which is neither of the counts this corpus separates — confirming that 2–3 here and 2–4 in 4.2 are guide-local figures to be reproduced on the exam, not Anthropic guidance.
  2. Code Review Anthropic The shipped implementation of this section: parallel specialized agents, a verification step that "checks candidates against actual code behavior to filter out false positives", dedup, and inline posting — with a `REVIEW.md` injected into every review agent as the criteria artifact.
  3. Increase output consistency Anthropic The documented levers that do move consistency — exact output formats, response prefill, constraining with examples, retrieval grounding — which is the positive list behind 4.1's negative claim about "be conservative". It also states 4.2's central mechanism plainly: "Provide examples of your desired output. This is more effective than abstract instructions."
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • The importance of explicit criteria over vague instructions (e.g., "flag comments only when claimed behavior contradicts actual code behavior" vs "check that comments are accurate")
  • How general instructions like "be conservative" or "only report high-confidence findings" fail to improve precision compared to specific categorical criteria
  • The impact of false positive rates on developer trust: high false positive categories undermine confidence in accurate categories

Skills in

  • Writing specific review criteria that define which issues to report (bugs, security) versus skip (minor style, local patterns) rather than relying on confidence-based filtering
  • Temporarily disabling high false-positive categories to restore developer trust while improving prompts for those categories
  • Defining explicit severity criteria with concrete code examples for each severity level to achieve consistent classification
Back to top