Numbering is global: a source keeps the same number everywhere it is cited, so [16] is the same page whether you meet it in Domain 1 or Domain 5. The number is a label, not an address — every link into this page points at the source's own name, so a link you saved keeps working even if the numbering shifts.
The exam is written against the exam guide, not against these pages. Where they differ — and they do, because the API moves faster than the certification — answer from the guide. The differences are called out where they matter.
All 73 sources verified ·Official documentation
- Handling stop reasons Anthropic The complete set of `stop_reason` values and the required handling for each — including `pause_turn`, which must be continued rather than raised.
- How tool use works Anthropic The canonical agentic loop as Anthropic writes it — the source the four-step lifecycle here restates.
- Handle tool calls Anthropic The exact `tool_result` message contract, including the case where text placed before `tool_result` blocks ends the turn early and returns a 400.
- Run agents in parallel Anthropic Where the hub is structurally enforced — and the one supported topology where it is not: agent teams, whose members share a task list and message each other directly.
- Subagents in the SDK Anthropic The "What subagents inherit" table, and the note that the spawning tool was renamed from `Task` to `Agent` in Claude Code v2.1.63.
- Agent SDK reference — TypeScript Anthropic The authoritative `AgentDefinition` field list, which is where you confirm the three fields taught here are the tested three rather than the only three.
- Parallel tool use Anthropic How multiple `tool_use` blocks in one assistant message are executed and returned together — the API layer beneath concurrent subagent spawning.
- Work with sessions (Agent SDK) Anthropic The actual fork contract — `resume` plus `forkSession` — and the warning that forking branches history, not the filesystem.
- Hooks reference Anthropic The `PreToolUse` schema that implements a prerequisite gate, and `PostToolUse`’s `updatedToolOutput`, which replaces a tool result before it reaches the model.
- Configure permissions Anthropic The allow/ask/deny rule layer and the `canUseTool` callback — the availability controls, none of which expresses an ordering constraint.
- Automate actions with hooks Anthropic Working shell-hook examples for validating commands and enforcing project rules — the shortest path from the argument here to something runnable.
- Intercept and control agent behavior with hooks Anthropic The SDK callback form of tool-call interception, for building on the Agent SDK rather than settings files.
- Orchestrate subagents at scale with dynamic workflows Anthropic What the three phases look like when someone has to run them on a codebase-wide audit or a 500-file migration.
- Manage sessions Anthropic The full resume surface — by name, by id, and the interactive picker — plus the caveat that an auto-generated title is not a resume handle.
- Checkpointing Anthropic Forking branches conversation history but not the filesystem; checkpointing is what rewinds file edits — the missing half of "no branch sees another’s work".
- Define tools Anthropic The normative tool-definition reference — name, `description`, `input_schema` — with Anthropic's own guidance on description quality, and the place to check whether the unit's four description components are the source's four. It is also where `tool_choice` is enumerated: "there are four possible options", `auto`, `any`, `tool` and `none`. With `any` or `tool` the API prefills the assistant message, so no natural-language text precedes the `tool_use` block.
- Troubleshooting tool use Anthropic The documented diagnostic order for wrong-tool-selected symptoms — where a reader checks whether auditing the system prompt for keyword-sensitive wording really is the recommended first move.
- Connect to external tools with MCP Anthropic The SDK's error-handling section, including per-server status reporting (`mcpServerStatus()` / `get_mcp_status()`) — the difference between a tool that failed and a server that never connected, which the four error categories do not cover.
- Manage tool context Anthropic The four documented remedies for exactly the pressure the 4–5 versus 18 rule describes: tool search, programmatic tool calling, prompt caching, and context editing. It answers the question the rule raises and closes off — what to do when you genuinely need 18 tools.
- Scale to many tools with tool search Anthropic On-demand tool loading at the Claude Code and Agent SDK layer — how the platform now expresses this unit's principle: keep few tools *in context*, not few tools *configured*.
- Connect Claude Code to tools via MCP Anthropic Documents three MCP installation scopes, not two, with a precedence order — local, then project, then user, then plugin servers, then claude.ai connectors — where local is the default and "the entire server entry from that source is used; fields are not merged across scopes". Local scope's stated purpose, "personal development servers, experimental configurations, or servers with credentials you don't want in version control", is what the guide assigns to user scope.
- Connect to MCP servers Anthropic The shortest path from "a tool is missing" to a fix: add a server, verify the connection with `/mcp`, and find the configuration on disk.
- Tools reference Anthropic The authoritative `Edit` contract, which names the two remedies Claude Code actually uses for a non-unique `old_string`: a longer string with surrounding context, or `replace_all: true`. The guide credits `Read` plus `Write` instead, and treats widening the anchor as a distractor.
- Common workflows Anthropic The step-by-step walkthrough of exploring an unfamiliar codebase — the incremental Grep-then-Read rule shown as an actual session.
- Set up Claude Code in a monorepo or large codebase Anthropic What changes at monorepo scale — code intelligence, sparse worktrees, per-package configuration — which is where the two-pass grep stops being sufficient.
- How Claude remembers your project Anthropic Settles the precedence question the guide leaves open: discovered CLAUDE.md files are concatenated in load order rather than overriding each other. Also adds the two locations the three-level table omits — machine-wide managed policy, which individual settings cannot exclude, and gitignored `CLAUDE.local.md` — and states that path-scoped rules trigger when Claude *reads* a matching file, that `@path` imports load at launch so they do not reduce context, and that rules with `paths:` are not re-injected after `/compact`.
- Debug your configuration Anthropic The diagnostic split behind "run `/memory` first": `/context` reports what actually loaded into the session, while `/memory` lists and opens the memory files for editing.
- Explore the .claude directory Anthropic The annotated map of everything Claude Code reads under `.claude/`, which shows `@import` targets, `rules/` with `paths:` globs and directory-level files as entries in one documented layout rather than isolated mechanisms.
- Extend Claude with skills Anthropic The frontmatter reference. `allowed-tools` grants pre-approval for the invoking turn and "does not restrict which tools are available: every tool remains callable" — the restricting field is `disallowed-tools`. It also documents the companions to `context: fork`: `agent` selects the subagent type, and `background` defaults to `true`, so the forked skill runs while you keep working.
- Agent Skills in the SDK Anthropic How the same `.claude/skills/` and `~/.claude/skills/` files load when the surface is the Agent SDK rather than the CLI — the case Scenario 4 keeps putting the reader in.
- Choose a permission mode Anthropic What plan mode concretely is: Claude reads files and runs shell commands to explore and writes a plan, but does not edit source, and edits stay blocked until you approve. Also the four ways in — `Shift+Tab`, `/plan`, `--permission-mode plan`, and `defaultMode`.
- Create custom subagents Anthropic Explore and Plan documented as built-in subagents, including the consequence the row here does not mention: both skip your CLAUDE.md files and the parent session’s git status to stay fast and cheap.
- Best practices for Claude Code Anthropic Anthropic’s own explore → plan → code sequence, which is what "the guide’s own combination" echoes, plus the test-first loop and the interview pattern as workflows performed on Claude Code itself.
- Prompting best practices Anthropic The consolidated reference on examples, clarity and structure. It recommends 3–5 examples for best results, which is neither of the counts this corpus separates — confirming that 2–3 here and 2–4 in 4.2 are guide-local figures to be reproduced on the exam, not Anthropic guidance.
- Run Claude Code programmatically Anthropic The normative page for non-interactive runs, including the flags that cannot combine with `-p` — `--bg`, and `--cloud` with a task description — a failure a CI author hits immediately.
- CLI reference Anthropic `--output-format` is a selector with three values — `text` (the default), `json`, `stream-json` — not a switch. It also settles the argument form this page hedges on: `--json-schema` takes the schema itself as an inline JSON string, is print-mode only, and exits with an error on an invalid schema.
- Get structured output from agents Anthropic Where the validated result actually lands — the `structured_output` field on the result message — plus the JSON Schema, Zod and Pydantic paths a pipeline step parses.
- Code Review Anthropic The shipped implementation of this section: parallel specialized agents, a verification step that "checks candidates against actual code behavior to filter out false positives", dedup, and inline posting — with a `REVIEW.md` injected into every review agent as the criteria artifact.
- Increase output consistency Anthropic The documented levers that do move consistency — exact output formats, response prefill, constraining with examples, retrieval grounding — which is the positive list behind 4.1's negative claim about "be conservative". It also states 4.2's central mechanism plainly: "Provide examples of your desired output. This is more effective than abstract instructions."
- Structured outputs Anthropic Documents a first-class structured-output path — `output_config.format`, formerly the `output_format` beta — that returns validated JSON directly without the extraction-tool trick, and explains how it differs from strict tool use. It also carries the supported Pydantic integration (`client.messages.parse()` with `output_format=Model`), which shows where schema validation stops and 4.4's semantic validator must begin.
- Strict tool use Anthropic The actual feature: `strict: true` on a tool definition, enforced by grammar-constrained sampling. It requires `additionalProperties: false` and `required`, and it drops several JSON Schema keywords — recursive `$ref`, `minimum`, `maximum`, `multipleOf`, `minLength`, `maxLength`.
- Reduce hallucinations Anthropic Anthropic's guidance on giving the model a legal way to say "not present", which is the general principle the nullable-field rule implements at the schema layer.
- Batch processing Anthropic 24 hours is an expiry, not a completion time: "Batches expire if processing does not complete within 24 hours", and an expired request returns no result and is not billed. The same page settles the tool question — "Tool use, including all server tools" and "Multi-turn conversations" are listed under What can be batched, and the real exclusion list is `stream`, `speed`, `store`, `previous_thread_event_id`, `cache_hint`, `context_hint`, `max_tokens: 0`.
- Retrieve Message Batch results Anthropic The API reference in normative form — "Batch results can be returned in any order… always use the `custom_id` field" — plus the `succeeded` / `errored` / `canceled` / `expired` result types this unit's resubmission logic has to branch on.
- Create a Message Batch Anthropic The actual `custom_id` constraint — 1–64 characters matching `^[a-zA-Z0-9_-]{1,64}$` — which rules out the document paths and URLs a reader would naturally reach for as join keys.
- Find bugs with ultrareview Anthropic The shipped instance of the independent-review architecture — a "multi-agent fleet with independent verification" where "every reported finding is independently reproduced and verified" — this unit's claim as a product decision rather than an assertion.
- Context windows Anthropic The normative statement of API statelessness and how the window accumulates across turns, which is the fact every trimming decision in 5.1 depends on.
- Context editing Anthropic The API-native alternative to hand-rolled trimming — server-side tool-result clearing with `clear_tool_uses_20250919` — for readers on the Messages API, where a `PostToolUse` hook does not exist.
- Customer support agent Anthropic The only Anthropic page treating escalation as designed behavior. Supplies the missing metric — escalation accuracy, target 95% — and places sentiment among business-impact metrics, never among routing signals.
- Handle approvals and user input Anthropic The SDK mechanism behind "ask for another identifier": `AskUserQuestion` and the `canUseTool` callback. Notes that `AskUserQuestion` is unavailable inside subagents spawned via the Agent tool.
- Explore the context window Anthropic An interactive simulation of what fills the window during a real session — what loads automatically, what each file read costs, what survives compaction.
- Commands Anthropic The reference entry for `/compact` and its focus instructions, plus the adjacent commands: `/context` to see what is filling the window, `/clear` when compaction is the wrong tool.
- Define success criteria and build evaluations Anthropic How to build the labeled evaluation set calibration presupposes, including deliberate edge-case coverage of "ambiguous test cases where even humans would find it hard to reach an assessment consensus".
- Citations Anthropic Model-generated citations bound to spans of a supplied document, returned as structured blocks — the artifact that makes an ambiguity flag checkable rather than a judgment call.
Specifications
- Tools Model Context Protocol The normative source for `isError`. It splits failures into protocol errors (JSON-RPC) and tool execution errors (`isError: true`), and states the reason this unit's whole argument depends on: "Clients SHOULD provide tool execution errors to language models to enable self-correction." It also fixes the result shape — model-visible data travels in `content` or `structuredContent`, and a tool returning structured content "SHOULD also return the serialized JSON in a TextContent block".
- Client Best Practices Model Context Protocol The spec's own quantified threshold for "too many tools": switch to progressive discovery once tool definitions occupy 1–5% of the context window. It turns the unit's unthresholded scoping rule into a number.
- Resources Model Context Protocol The normative resource model — URI-identified, listable, subscribable — which is what makes "read-the-index instead of search-to-discover" mechanically true rather than a metaphor.
- NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0) NIST A federal framework stating the segment rule directly: "Accuracy measurements may include disaggregation of results for different data segments", paired with representative test sets and human intervention where the system cannot correct itself.
- Survey Methods and Practices (Catalogue no. 12-587-X) Statistics Canada The national statistical agency's survey-design handbook — the standard reference for stratified sampling, covering sample size, allocation across strata, and selection.
- Section 2. Stratified sampling (Survey Methodology 45(2), Cat. 12-001-X) Statistics Canada The load-bearing property, stated exactly: stratification "enables controlling sample sizes and precision of estimates for the strata", and can improve overall precision for a fixed cost.
Anthropic Engineering
- How we built our multi-agent research system Anthropic Engineering, 13 Jun 2025 The field report where narrow decomposition was observed and measured — one subagent on the 2021 chip crisis while two duplicated work on 2025 supply chains.
- Building effective AI agents Anthropic Engineering, 19 Dec 2024 The source of the vocabulary, with all five named patterns and the workflow-versus-agent distinction that the two-row table here is a projection of.
- Effective context engineering for AI agents Anthropic Engineering, 29 Sep 2025 Names and explains context rot — recall degrades as the window fills — which is the evidence that a bigger window buys capacity, not attention.
- Writing effective tools for agents — with agents Anthropic Engineering The `analyze_content` versus `analyze_document` problem treated at length, with the opposite default: namespacing (`asana_search` vs `jira_search`) and consolidation over splitting. That is the tension this unit resolves by assertion when it says "fewer overlapping, more purpose-specific".
- How we built Claude Code auto mode: a safer way to skip permissions Anthropic Engineering Anthropic's own shipped escalation trigger: three consecutive or twenty total denials stop the model and escalate to a human. A counted, externally observable signal rather than a self-reported score.
Research
- Language Models (Mostly) Know What They Know Kadavath et al., Anthropic (2022) Larger models are well calibrated on multiple-choice and true/false questions given the right format, and `P(True)` self-evaluation improves with scale — but models struggle with calibration of P(IK) on new tasks.
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs Xiong et al., ICLR 2024 LLMs verbalizing confidence "tend to be overconfident, potentially imitating human patterns", and every elicitation method evaluated struggles on challenging tasks.
- Teaching Models to Express Their Uncertainty in Words Lin, Hilton & Evans, TMLR 2022 Verbalized confidence maps to well-calibrated probabilities and stays moderately calibrated under distribution shift, grounded in latent representations that correlate with epistemic uncertainty.
- How angry are your customers? Sentiment analysis of support tickets that escalate Werner et al., AffectRE @ RE 2018 The one empirical study on the question: a considerable sentiment difference between escalated and non-escalated support tickets. Association with escalation, not with complexity, and the authors call it preliminary.
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification Buolamwini & Gebru, PMLR 81 (FAT* 2018) The canonical measured instance of aggregate masking: 0.8% error for lighter-skinned males against 34.7% for darker-skinned females, in systems whose headline accuracy looked acceptable.
- Model Cards for Model Reporting Mitchell et al., FAT* 2019 Establishes disaggregated evaluation as a reporting norm rather than an optional analysis: benchmarked performance broken out by group and by intersection of groups.
- On Calibration of Modern Neural Networks Guo, Pleiss, Sun & Weinberger, ICML 2017 Raw model confidence is not trustworthy as a probability, and temperature scaling is the post-hoc procedure that fixes it against held-out labeled data. Measured on image and document classifiers, not on verbalized LLM self-report.
- Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer Madras, Pitassi & Zemel, NeurIPS 2018 The literature on when a model should hand off to a human. Supports "one axis is insufficient"; its second axis is the human reviewer's competence, not document ambiguity.