Note on selection
Design and implement agentic loops for autonomous task execution
What you need to know
-
stop_reasondrives termination. The guide tests two values:"tool_use"iterates,"end_turn"exits. - Push back the whole assistant message, then all matching
tool_resultblocks in one user message. - Nothing is discarded between turns; the growing history is the reasoning substrate.
- Model-driven means Claude composes the sequence at runtime. A decision tree ships only your predictions.
- A tripped iteration cap is an incident to alert on, never a finished task.
An agentic loop is the smallest unit of autonomy in the Claude Agent SDK. You hand Claude a goal and a tool inventory. Claude — not your code — picks the next tool from everything it has learned so far.
The lifecycle
Each iteration has four steps:
- Send the current conversation (system prompt, tools, full message history) to the model.
- Inspect
stop_reason, and branch on it alone. - Execute every requested tool. One assistant message may carry several
tool_useblocks — run them and collect all results. - Append and repeat. Push the assistant message in full (the
tool_useblocks must survive), then push one user message containing every matchingtool_result.3 Loop back to step 1.
stop_reason |
Claude's meaning | Loop action |
|---|---|---|
"tool_use" |
Needs tools run before it can continue | Execute, append results, iterate |
"end_turn" |
Turn finished, answer produced | Exit — the only normal exit |
Any other value — the guide names max_tokens; the live API returns more |
Not a completion | Branch explicitly; never treat as "complete" |
Those are the two values the certification grades. The live API returns more — stop_sequence, pause_turn and refusal among them — so a production loop branches on the value rather than bucketing everything else as an error.1
Two invariants make this work.2 First, stop_reason is a structured field the model emits deliberately, so branching on it is deterministic. Second, tool results are appended to history, not consumed and discarded. That accumulation is what lets Claude reason about the next action: having seen lookup_order return "delivered", it can decide a replacement is appropriate without you writing a rule for it.
Model-driven vs pre-configured
A pre-configured decision tree or fixed tool sequence encodes the paths you anticipated. A model-driven loop encodes the goal and lets Claude compose paths at runtime — including ones you never enumerated. That is the whole point of the pattern, and it is why "add a routing classifier that pre-selects the tool" is almost always the wrong answer on the exam: it replaces reasoning with keyword matching.
The trade is that termination must be read from the protocol, not guessed from the prose. Claude often writes a friendly closing sentence while still requesting a tool, so scanning for "done" or "I have completed" ends turns early and mid-task. A hard cap used as the primary stopping mechanism truncates legitimate long investigations. Caps and timeouts are circuit breakers for runaway behavior — logged and alerted on — never the normal exit path.
Walk the four steps of the loop — send the request, inspect stop_reason, execute the requested tools, append their results — and show that the only transition signal is the structured stop_reason field: "tool_use" sends tool results back into conversation history before the next request, "end_turn" leaves the loop, and an iteration or time limit is the other exit.
Click a transition to highlight which stop_reason value produces it.
Worked examples
The canonical loop, written correctly
Scenario 1 · Customer Support Resolution AgentA support agent has get_customer, lookup_order, process_refund and escalate_to_human exposed as MCP tools. The loop below never inspects text and never counts iterations to decide completion — it branches on stop_reason alone, and it returns all tool results in one user message so parallel tool use keeps working.
const messages: MessageParam[] = [{ role: 'user', content: userRequest }];
while (true) {
const response = await client.messages.create({
model: 'claude-opus-5',
max_tokens: 4096,
tools,
messages,
});
// Append the FULL assistant content — tool_use blocks must be preserved.
messages.push({ role: 'assistant', content: response.content });
if (response.stop_reason === 'end_turn') break; // the only normal exit
if (response.stop_reason === 'tool_use') {
const toolUses = response.content.filter((b) => b.type === 'tool_use');
const results = await Promise.all(
toolUses.map(async (b) => ({
type: 'tool_result' as const,
tool_use_id: b.id,
content: await executeTool(b.name, b.input),
})),
);
// ALL results go back in ONE user message; this becomes the next iteration's context.
messages.push({ role: 'user', content: results });
continue;
}
// The guide names only max_tokens here. The live API also returns
// stop_sequence, pause_turn and refusal — and pause_turn is resumable, not
// fatal: a server-side tool hit its iteration limit, so re-send with the
// assistant turn appended rather than raising.
if (response.stop_reason === 'pause_turn') continue;
throw new AgentLoopError(response.stop_reason);
}Emergent tool sequences you never wrote down
Scenario 1 · Customer Support Resolution AgentA customer writes: "my order arrived smashed, and I think I was double-charged last month." No decision tree in the codebase covers "damaged goods plus a billing dispute". In a model-driven loop, Claude calls get_customer to establish identity, then lookup_order for the damaged item, notices the second concern in the accumulated context, looks up the billing history, and only then decides whether process_refund covers both or whether the policy exception requires escalate_to_human.
The correct architectural response to a case like this is to keep the loop model-driven and improve the inputs to reasoning — richer tool descriptions, clearer escalation criteria — not to add a pre-turn classifier that picks tools by keyword. A classifier would have routed on "smashed" and silently dropped the billing concern.
Distinguishing a circuit breaker from a stopping mechanism
Scenario 4 · Developer Productivity with ClaudeA codebase-exploration agent is given "find every place we validate email addresses". Legitimate runs take 30–60 tool calls across Grep, Glob and Read. A team ships if (iterations > 10) return partialAnswer and immediately sees truncated, confidently-wrong answers.
The fix is to separate the two concerns. Termination stays bound to stop_reason === "end_turn". A cap of, say, 200 iterations remains — but it throws, logs, and alerts, because reaching it means the agent is looping, not that it finished. Same for a wall-clock timeout: it is an incident signal, not a result.
const MAX_ITERATIONS = 200; // runaway protection only
let iterations = 0;
while (true) {
if (++iterations > MAX_ITERATIONS) {
// Loud failure — NOT a successful completion.
metrics.increment('agent.runaway');
throw new RunawayLoopError({ iterations, lastToolCalls: recentToolNames });
}
const response = await step(messages);
if (response.stop_reason === 'end_turn') return response; // the real exit
// ... execute tools, append results, continue
}Anti-patterns
- Parsing natural-language signals ("I'm done", "that completes the task") from the assistant text to terminate the loop, instead of branching on
stop_reasonbecause Claude routinely writes conversational prose in the same message as atool_useblock and the loop exits mid-task. - Using an arbitrary iteration cap as the primary stopping mechanism instead of
stop_reason === "end_turn"because legitimate investigations vary enormously in length and the cap silently truncates them into confident partial answers. - Treating the presence of assistant text content as a completion indicator instead of reading
stop_reasonbecause text and tool requests coexist in the same response. - Executing tools but not appending their results to conversation history (or appending only a summary of them) instead of returning proper
tool_resultblocks because the model then re-requests the same tools or reasons from stale context.
How it is examined
- When a stem describes a loop that "sometimes stops early" or "returns incomplete answers", scan the options for the one that binds termination to
stop_reason; options mentioning text parsing, keyword detection, or iteration limits are the distractors. - Options that add a pre-turn routing classifier, a keyword-based tool selector, or a fixed tool sequence are testing whether you understand model-driven decision-making — they are almost always over-engineered and wrong unless the stem explicitly requires deterministic ordering (that is 1.4 territory).
- Watch for options that "summarize tool results to save tokens" before appending them; discarding raw tool results breaks the reasoning chain this statement is about.
Beyond the exam — what the API does that this does not grade
The rest of stop_reason. The guide tests "tool_use" and "end_turn". The live Messages API also returns:1
| Value | What it means | Loop action |
|---|---|---|
max_tokens |
Output cap hit mid-answer | Raise max_tokens, or stream |
stop_sequence |
A custom stop sequence matched | Treat as your own protocol, not completion |
pause_turn |
A server-side tool hit its internal iteration limit | Re-send the conversation with the assistant turn appended. Do not add a "continue" message and do not throw |
refusal |
Safety classifiers declined | Read stop_details.category; content may be empty. Do not retry unchanged |
stop_details is populated only when stop_reason is "refusal" — guard before reading it.
References — 3 sources
- Handling stop reasons Anthropic The complete set of `stop_reason` values and the required handling for each — including `pause_turn`, which must be continued rather than raised.
- How tool use works Anthropic The canonical agentic loop as Anthropic writes it — the source the four-step lifecycle here restates.
- Handle tool calls Anthropic The exact `tool_result` message contract, including the case where text placed before `tool_result` blocks ends the turn early and returns a 400.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- The agentic loop lifecycle: sending requests to Claude, inspecting stop_reason ("tool_use" vs "end_turn"), executing requested tools, and returning results for the next iteration
- How tool results are appended to conversation history so the model can reason about the next action
- The distinction between model-driven decision-making (Claude reasons about which tool to call next based on context) and pre-configured decision trees or tool sequences
Skills in
- Implementing agentic loop control flow that continues when stop_reason is "tool_use" and terminates when stop_reason is "end_turn"
- Adding tool results to conversation context between iterations so the model can incorporate new information into its reasoning
- Avoiding anti-patterns such as parsing natural language signals to determine loop termination, setting arbitrary iteration caps as the primary stopping mechanism, or checking for assistant text content as a completion indicator