Domain 5 · 15% of the exam
Context Management & Reliability
Domain 5 measures whether you can keep an agent correct over time: preserving critical facts across long conversations and compaction steps, escalating instead of guessing when a request is ambiguous or out of policy, propagating failures across a multi-agent system so a coordinator can recover, exploring large codebases without context degradation, routing work to human reviewers with calibrated confidence, and preserving provenance when findings from many sources are merged. It carries the smallest blueprint weight (15%), but it is a primary domain in four of the six exam scenarios — Customer Support (1), Code Generation with Claude Code (2), Multi-Agent Research (3) and Structured Data Extraction (6) — so its items are spread across most of the exam rather than concentrated in one block. The recurring theme is that reliability comes from what you deliberately preserve, structure and annotate, not from asking the model to be more careful.
6 statements 18 examples 11 diagrams ~18 min
Take the quiz Study map
Choose a task statement
Each statement is a focused page. Your progress and quiz links still connect through the domain.
- 5.1 1 of 6
Manage conversation context to preserve critical information across long interactions
Pin the numbers outside the summary: a case-facts block, re-sent verbatim, one structured layer per issue. ~2 min · 3 examples · 2 diagrams - 5.2 2 of 6
Design effective escalation and ambiguity resolution patterns
Valid triggers: an explicit request for a human, a policy exception or policy gap, and inability to make meaningful progress — not "the case looks complex". ~5 min · 3 examples · 1 diagrams - 5.3 3 of 6
Implement error propagation strategies across multi-agent systems
A propagated error must carry failure type, what was attempted, partial results and suggested alternatives — that is what makes coordinator recovery… ~3 min · 3 examples · 2 diagrams - 5.4 4 of 6
Manage context effectively in large codebase exploration
The tell-tale sign of context degradation is answers drifting from specific discovered classes to generic "typical patterns", plus inconsistency between… ~2 min · 3 examples · 2 diagrams - 5.5 5 of 6
Design human review workflows and confidence calibration
A 97% aggregate can hide a document type or a field performing far worse; validate accuracy by document type and by field before reducing review. ~4 min · 3 examples · 2 diagrams - 5.6 6 of 6
Preserve information provenance and handle uncertainty in multi- source synthesis
Claim-source mappings (source URL or document name plus relevant excerpt) must be created at discovery and preserved and merged through synthesis, never… ~2 min · 3 examples · 2 diagrams