Skip to content
CCAR-FAcademy
Domain 3 · Statement 3.5 5 of 6
3.5

Apply iterative refinement techniques for progressive improvement

  • Concrete input/output examples are the most effective fix when prose descriptions are interpreted inconsistently — 2 to 3 pairs is the sweet spot (distinct from the 2 to 4 few-shot examples of Domain 4 task statement 4.2).
  • Test-driven iteration means writing tests for behavior, edge cases and performance first, then iterating by sharing the actual failures.
  • The interview pattern has Claude ask questions before implementing, surfacing considerations the developer never thought to specify.
  • Interacting problems go in one detailed message; independent problems are better fixed sequentially.
  • To fix a stubborn edge case, supply the exact offending input and the exact expected output rather than describing the rule again.

Iterative refinement is the craft of turning a mediocre first result into a correct one with the fewest, best-targeted follow-up messages. The exam tests four specific techniques and, crucially, when each one applies.

Concrete input/output examples

When prose descriptions are interpreted inconsistently, concrete input/output examples are the most effective way to communicate the expected transformation. "Normalize the address field" can mean five things; two or three worked pairs — this input becomes exactly this output — pin the semantics down in a way no adjective can. The guide's guidance is 2–3 examples: enough to show the pattern and its edge, not so many that you are writing the test suite by hand.41 Keep that figure attached to its scope — 2–3 is the count for input/output pairs used in iterative refinement, and it is a different statement from Domain 4 §4.2's 2–4 targeted few-shot examples for ambiguous-case prompting. A stem asking "how many examples?" is answered by which of the two techniques it describes. This is also the fix for a specific edge-case failure: supply a concrete input containing the problem value (a null, an empty string, a date at a boundary) alongside the exact expected output.

Test-driven iteration

Write the test suite first, covering expected behavior, edge cases and performance requirements, then let Claude implement against it and iterate by sharing the failures. Each failure is unambiguous, machine-generated feedback: no interpretation gap, no argument about whether the behavior is correct. Progressive improvement is then just a loop — run tests, paste failures, repeat — and the loop has an objective stopping condition.

The interview pattern

In an unfamiliar domain, you do not know what you failed to specify. The interview pattern inverts the flow: ask Claude to ask you questions before implementing. Prompting for questions about a caching layer surfaces invalidation strategy, stampede behavior, and what happens when the cache is unavailable — considerations the developer may not have anticipated. It converts unknown unknowns into an answerable list, before code exists that assumes the wrong answers.40

Batch or sequential?

The judgment call the exam likes most. Give all issues in a single detailed message when the problems interact — if fixing the retry logic changes how the timeout should behave, fixing them one at a time makes each fix invalidate the last. Fix sequentially when the problems are independent, because smaller focused messages produce more reliable, easier-to-verify changes and you can stop early if one fix reveals something new.

The unifying principle: reduce ambiguity per round trip. Examples, tests, and questions are all ways of making the next iteration's feedback specific.

Picking the right refinement techniqueMap the symptom you are seeing to the refinement technique the guide prescribes, and show test-driven iteration as a closed loop with an objective stopping condition.First result is unsatisfactoryFirst result isunsatisfactoryOutput varies between runs?Outputvariesbetweenruns?Give 2-3 input/output examplesGive 2-3input/outputexamplesUnfamiliar domain, unclear requirements?Unfamiliardomain,unclearrequirements?Interview pattern (Claude asks questions)Interviewpattern (Claudeasksquestions)Write tests firstWrite testsfirstShare test failuresShare testfailuresAll tests pass — doneAll testspass — donediagnose thesymptomyes, pin thetransformationnoyes, surfaceconsiderations firstno, requirementsalready clearimplement and runiterate on failuresonlyobjective stoppingcondition
Picking the right refinement technique

Map the symptom you are seeing to the refinement technique the guide prescribes, and show test-driven iteration as a closed loop with an objective stopping condition.

one transition at a time
One message or several? Interacting vs independent issuesMake the batching decision concrete: interacting fixes must be described together, independent fixes are more reliably handled one at a time.Several issues found in reviewSeveralissues foundin reviewDo fixes affect each other?Do fixesaffect eachother?Interacting issuesInteractingissuesIndependent issuesIndependentissuesOne message with all issuesOnemessagewith allissuesSequential focused iterationsSequentialfocusediterationsCoherent combined behaviorCoherentcombinedbehaviorEach change easy to verifyEachchangeeasy toverifythe key questionyes, e.g. TTL andinvalidationno, e.g. log text anda typofix togetherno fix invalidatesanotherfix one at a timesmaller blastradius
One message or several? Interacting vs independent issues

Make the batching decision concrete: interacting fixes must be described together, independent fixes are more reliably handled one at a time.

one transition at a time

Two examples beat three paragraphs of prose

Scenario 2 · Code Generation with Claude Code

A migration script must normalize legacy customer records. Prose instructions ("clean up the phone numbers, keep the country code if present, drop extensions") produce a different interpretation on every run.

Replacing the prose with concrete pairs ends the ambiguity, and deliberately including a null and an empty-extension case fixes the edge-case handling that prose kept getting wrong.

text
Transform each record's phone field. Examples:

  in:  { phone: "+34 91 555 0123 ext. 22" }
  out: { phone: "+34915550123", extension: "22" }

  in:  { phone: "915550123" }
  out: { phone: "+34915550123", extension: null }

  in:  { phone: null }
  out: { phone: null, extension: null }      # do not throw, do not default

Anything not matching these shapes: leave the record untouched and add its
id to the skipped list.
2-3 worked pairs, including the edge cases that keep failing

Test-driven iteration on a rate limiter

Scenario 4 · Developer Productivity with Claude

Rather than describing a token-bucket rate limiter in prose, the engineer writes the suite first: normal traffic passes, burst above capacity is rejected, the bucket refills over time, concurrent callers are accounted correctly, and 10k checks complete inside the performance budget.

Claude implements against the suite; the engineer runs it and pastes the failures verbatim. Round one fails the refill test, round two fails the concurrency test, round three passes. Each iteration is driven by unambiguous evidence instead of a judgment call, and "done" is defined before the work starts.

python
def test_allows_traffic_within_capacity(): ...
def test_rejects_burst_above_capacity(): ...
def test_refills_tokens_over_time(): ...        # edge case
def test_concurrent_callers_share_one_bucket(): # edge case
    ...
def test_10k_checks_under_50ms():               # performance requirement
    ...

# Then: implement, run, and paste the failing output back verbatim
# as the next iteration's input.
The suite is written first — expected behavior, edge cases, performance

Interview first, then batch the interacting fixes

Scenario 4 · Developer Productivity with Claude

An engineer must add a caching layer to a service in a domain they do not know well. Before any code, they ask Claude to interview them. The questions surface invalidation strategy, behavior on cache unavailability, and whether stale reads are acceptable during a write — none of which were in the original request.

After implementation, review finds four issues. Three are independent (a misleading log message, a missing metric, a typo in a config key) and are handled one at a time. The other two — the TTL and the invalidation-on-write path — interact: changing one without the other produces incoherent behavior, so they go into a single detailed message describing both, together with the intended combined semantics.

  • Rewriting the prose description a third time instead of supplying 2-3 concrete input/output examples because prose is exactly what was being interpreted inconsistently.
  • Implementing first and writing tests afterwards instead of test-driven iteration because you lose the objective failure signal that drives progressive improvement.
  • Fixing interacting issues one at a time instead of in a single detailed message because each isolated fix invalidates the assumptions of the previous one.
  • Dumping every unrelated independent issue into one giant message because the changes become hard to verify and a failure in one obscures the others.
  • If the stem says results are "inconsistent" or "interpreted differently each time", the answer is concrete input/output examples — not a longer or more emphatic description.
  • Watch the interaction test: options will offer both "send all issues at once" and "fix them one by one". Decide by whether the fixes affect each other, and re-read the stem for that hint.
  • The interview pattern shows up whenever the developer is in an unfamiliar domain; the distractor is asking Claude to "explain its approach" afterwards, which surfaces nothing new before the code exists.
References — 2 sources
  1. Best practices for Claude Code Anthropic Anthropic’s own explore → plan → code sequence, which is what "the guide’s own combination" echoes, plus the test-first loop and the interview pattern as workflows performed on Claude Code itself.
  2. Prompting best practices Anthropic The consolidated reference on examples, clarity and structure. It recommends 3–5 examples for best results, which is neither of the counts this corpus separates — confirming that 2–3 here and 2–4 in 4.2 are guide-local figures to be reproduced on the exam, not Anthropic guidance.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • Concrete input/output examples as the most effective way to communicate expected transformations when prose descriptions are interpreted inconsistently
  • Test-driven iteration: writing test suites first, then iterating by sharing test failures to guide progressive improvement
  • The interview pattern: having Claude ask questions to surface considerations the developer may not have anticipated before implementing
  • When to provide all issues in a single message (interacting problems) versus fixing them sequentially (independent problems)

Skills in

  • Providing 2-3 concrete input/output examples to clarify transformation requirements when natural language descriptions produce inconsistent results
  • Writing test suites covering expected behavior, edge cases, and performance requirements before implementation, then iterating by sharing test failures
  • Using the interview pattern to surface design considerations (e.g., cache invalidation strategies, failure modes) before implementing solutions in unfamiliar domains
  • Providing specific test cases with example input and expected output to fix edge case handling (e.g., null values in migration scripts)
  • Addressing multiple interacting issues in a single detailed message when fixes interact, versus sequential iteration for independent issues
Back to top