Note on selection
Design multi-instance and multi-pass review architectures
What you need to know
- A model that just generated code retains the reasoning behind it, so it is unlikely to question its own decisions in the same session.
- An independent instance with no prior reasoning context catches subtle issues better than a self-review instruction or extended thinking.
- Split large multi-file reviews into per-file passes for local issues plus separate integration passes for cross-file data flow.
- Merge the two finding sets rather than concatenating them: dedupe on (location, detected_pattern), integration passes win on contract questions and local passes win on line-level ones. Why a single giant pass degrades is covered in 1.6.
- A verification pass where the model self-reports confidence per finding enables calibrated routing (auto-comment vs human triage).
- Confidence is routing metadata attached to findings, not the criterion that decides what counts as a finding (see 4.1).
- Routing thresholds must be calibrated against a labeled validation set (5.5), and confidence is never an escalation trigger in a live conversation (5.2) — the axis is three-way: banned upstream (4.1), credited as calibrated downstream routing (4.6), rejected in-flight (5.2).
Why a model is a poor reviewer of its own work
When a model generates code and is then asked, in the same session, to review it, it still holds the reasoning that produced that code. Every design decision already looks justified, because the justification is right there in context. The result is that self-review is less likely to question its own decisions — it tends to confirm rather than challenge. This is a context property, not a capability shortfall. That is why the guide is explicit that independent review instances, running without the generator's prior reasoning context, catch subtle issues more effectively than either a self-review instruction ("now critically review your work") or extended thinking.53 Adding more reasoning inside the same context does not remove the bias built into that context.
Practically: for generated code, spin up a second Claude instance that receives the code and the requirements but not the generation transcript.15 It has to reconstruct intent from the artifact, which is exactly what a human reviewer does and exactly what surfaces the assumption the generator never questioned.
Multi-pass review: the pass taxonomy and the merge rule
Why a single giant pass degrades is covered in full in 1.6 — attention dilution, contradictory findings inside one PR, and why a larger context window is not the remedy. Here the subject is the shape of the decomposition and what becomes of the findings afterwards.
The taxonomy has exactly two pass types, each defined by the question it can answer.
| Pass type | The question it answers | What it structurally cannot answer | Wins the merge on |
|---|---|---|---|
| Per-file local pass — one per changed file, sees that file's diff | logic errors, error handling, unchecked inputs | whether callers satisfy this file's new preconditions | line-level questions — the only pass that saw the file in detail |
| Separate integration pass — one per data flow, sees the changed contracts and every call site touching them | cross-file data flow: does the emitted type match what the consumer expects, does the caller apply the discount the callee assumes, does a contract change break a distant call site | line-level detail inside unrelated files | contract and data-flow questions — the only pass that saw both sides |
Because the two types answer different questions they get different prompts, and they produce two finding sets that must be merged rather than concatenated.45 The merge rule falls straight out of the taxonomy:
- Dedupe on (location,
detected_pattern). - Resolve each collision in favor of the pass that could see the answer, as the last column above records.
- Pass through untouched any finding that only one pass type could have produced.
Verification passes with self-reported confidence
The third technique is a verification pass in which the model self-reports confidence alongside each finding, so findings can be routed by calibration. High-confidence findings post automatically as PR comments, low-confidence findings go to a human triage queue, and the rest are dropped or batched. Note the careful distinction against 4.1 — confidence is used downstream, as routing metadata on an already-generated finding, not upstream as the criterion that decides whether the model should report at all. Asking the model to filter by confidence during finding does not improve precision; asking it to label findings with confidence so your pipeline can route them does add value.
Two preconditions bound that value, and both live in Domain 5:
- The thresholds must be calibrated. A self-reported "high" is not a probability until you have measured what it means. The procedure is 5.5's: have the model emit field-level or finding-level confidence, then set the routing thresholds against a labeled validation set, so "auto-post above this level" reflects a measured accuracy rather than an assumption. An uncalibrated threshold is 4.1's anti-pattern wearing a number.
- Routing is not escalation. Confidence is never the trigger for handing a live conversation to a human (5.2). Official Sample Question 3 rejects precisely that option: the agent is already incorrectly confident on the hard cases, so a confidence threshold escalates the easy ones and keeps the hard ones. Explicit escalation criteria with few-shot examples are the credited fix there. Routing works here because a finding already exists and is being sorted offline against calibrated data; escalation fails because the decision must be made in-flight by the same miscalibrated model.
Show the two pass types with the question each asks and the question each structurally cannot answer, then how their findings are merged by authority and routed by calibrated self-reported confidence.
Hover a pass to see which question it asks and which it structurally cannot answer.
Worked examples
A second independent instance reviews the generated code
Scenario 5 · Claude Code for Continuous IntegrationThe key detail is the message history: the reviewer call starts from an empty conversation and receives the requirements plus the diff, never the generator's transcript or thinking. It must reconstruct intent from the artifact, which is what makes it able to disagree with the generator's assumptions.
// Pass 1 -- generation. This conversation holds all the design reasoning.
const generation = await client.messages.create({
model: 'claude-opus-5',
max_tokens: 8192,
messages: generatorHistory, // requirements, exploration, revisions
});
const code = extractCode(generation);
// Pass 2 -- review by an INDEPENDENT instance. Fresh message list:
// the generator's reasoning is deliberately NOT carried over.
const review = await client.messages.create({
model: 'claude-opus-5',
max_tokens: 8192,
system: reviewerCriteria, // explicit report/skip criteria (4.1)
tools: [reportFindings], // schema-enforced findings (4.3)
tool_choice: { type: 'tool', name: 'report_findings' },
messages: [{ role: 'user', content: renderReviewRequest(requirements, code) }],
});
// Anti-pattern for contrast:
// generatorHistory.push({ role: 'user', content: 'Now review your own code.' })
// -- same session, same reasoning context, confirms rather than challenges.Per-file passes plus a cross-file integration pass
Scenario 5 · Claude Code for Continuous IntegrationA 40-file PR is decomposed into the two pass types, each defined by the question it can answer and the question it structurally cannot, and the two finding sets are then reconciled by an explicit merge rule. Note that the merge is not a concatenation: findings are deduped and every collision is resolved in favor of whichever pass had the visibility to answer it.
LOCAL PASSES (one per changed file, N = 40)
Input: the single file's diff + surrounding context for that file
Asks: logic errors, unchecked inputs, error handling, resource leaks
Cannot: judge whether callers satisfy this file's new preconditions
INTEGRATION PASSES (grouped by data flow, N = 4)
Input: the signatures/contracts that changed + every call site touching them
Asks: type and contract mismatches across module boundaries,
state mutated in one module and read in another,
a precondition assumed by the callee and not met by the caller
Cannot: see line-level detail inside unrelated files
MERGE (a reconciliation, not a concatenation)
1. Dedupe on (location, detected_pattern).
2. On a collision, the pass that could see the answer wins:
contract / data-flow question -> integration finding
line-level question -> local finding
3. Findings only one pass type could produce pass through unchanged.
Result: one finding set in which no two entries disagree about the same
construct, because each question type has exactly one authoritative pass.
TAIL
A verification pass then attaches self-reported confidence to each merged
finding, and the pipeline routes on it -- thresholds calibrated against a
labeled validation set (5.5), never used to escalate a live conversation.Confidence-annotated findings for calibrated routing
Scenario 5 · Claude Code for Continuous IntegrationA verification pass re-examines the merged findings and attaches a self-reported confidence to each. The pipeline routes on that value: high confidence posts automatically, medium goes to a human triage queue, low is recorded for pattern analysis but not shown. Confidence is used to route — not to decide what the reviewer should have looked for.
[
{
"location": "src/billing/invoice.ts:88",
"issue": "total computed before discount is applied",
"severity": "HIGH",
"detected_pattern": "reduce_over_items_without_discount_call",
"confidence": "high",
"route": "post_as_pr_comment"
},
{
"location": "src/sync/worker.ts:142",
"issue": "possible race on the cursor when two workers start together",
"severity": "MEDIUM",
"detected_pattern": "shared_cursor_without_lock",
"confidence": "medium",
"route": "human_triage_queue"
}
]Anti-patterns
- In-session self-review: appending "now critically review your own code" to the generating conversation instead of spawning an independent instance because the model still holds the reasoning that produced the code and confirms rather than challenges it.
- Extended thinking as a substitute for independent review because more reasoning inside the same context does not remove the context bias that self-review suffers from.
- One-shot mega review: sending every changed file in a single review prompt instead of per-file passes plus separate integration passes because one pass collapses the two question types into one context and the local and cross-file findings compete there instead of being merged under an explicit rule (the degradation mechanism itself is covered in 1.6).
- Dropping confidence from findings and treating every finding as equal instead of having the verification pass self-report confidence because you lose the signal that would let you route findings to auto-comment or human triage.
How it is examined
- When a stem says an agent generates code, reviews it, and still misses subtle bugs, the credited answer is a second independent instance without the generator’s reasoning context. "Add a self-review instruction", "enable extended thinking", and "raise the reasoning effort" are the standard distractors.
- When a stem asks what happens after decomposition — how two sets of findings get reconciled, or which pass type owns a given kind of issue — the answer is the two-type taxonomy plus an explicit merge step (dedupe, integration wins on contracts, local wins on line detail). The prior question of why one giant pass degrades at all is 1.6 territory.
- Confidence is a three-way axis, so read the exact role the stem assigns it: (a) an upstream filter deciding what counts as a finding is a distractor (4.1); (b) downstream routing metadata on an already-generated finding, with thresholds calibrated on a labeled validation set, is credited (4.6 with 5.5); (c) a threshold that escalates a live conversation to a human is wrong (5.2, official Sample Question 3 — the agent is already incorrectly confident on the hard cases). Between two routing options, the one that mentions calibration is the stronger.
References — 3 sources
- Orchestrate subagents at scale with dynamic workflows Anthropic What the three phases look like when someone has to run them on a codebase-wide audit or a 500-file migration.
- Code Review Anthropic The shipped implementation of this section: parallel specialized agents, a verification step that "checks candidates against actual code behavior to filter out false positives", dedup, and inline posting — with a `REVIEW.md` injected into every review agent as the criteria artifact.
- Find bugs with ultrareview Anthropic The shipped instance of the independent-review architecture — a "multi-agent fleet with independent verification" where "every reported finding is independently reproduced and verified" — this unit's claim as a product decision rather than an assertion.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Self-review limitations: a model retains reasoning context from generation, making it less likely to question its own decisions in the same session
- Independent review instances (without prior reasoning context) are more effective at catching subtle issues than self-review instructions or extended thinking
- Multi-pass review: splitting large reviews into per-file local analysis passes plus cross-file integration passes to avoid attention dilution and contradictory findings
Skills in
- Using a second independent Claude instance to review generated code without the generator's reasoning context
- Splitting large multi-file reviews into focused per-file passes for local issues plus separate integration passes for cross-file data flow analysis
- Running verification passes where the model self-reports confidence alongside each finding to enable calibrated review routing