Skip to content
CCAR-FAcademy
Domain 5 · Statement 5.1 1 of 6
5.1

Manage conversation context to preserve critical information across long interactions

  • Pin the numbers outside the summary: a case-facts block, re-sent verbatim, one structured layer per issue.
  • The API stores nothing. Coherence is the full history you resend, so whatever you drop is unrecoverable.
  • Beginning and end survive best; recall decreases toward the middle as the window fills. Summary first, detail below under headers.
  • Trim at the tool boundary, not afterwards — 40 fields where 5 matter ride along on every future request.
  • Upstream agents ship facts, citations and scores; a small downstream budget cannot afford narrative.

Long sessions fail in predictable ways, and the exam expects the mechanism behind each failure rather than the vocabulary.

Summarization loses precision first. Compression favors narrative over data. "USD 847.32 charged on March 14 for order A-99231" becomes "a billing issue from last month"; "a refund within 5 business days" becomes "the customer expects resolution soon". Numerical values, percentages, dates and customer-stated expectations are what a summarizer treats as noise, and what the agent needs to act correctly. The fix is not a better summarizer. It is to keep transactional facts outside the summarized region, in a case facts block re-included in every prompt. A session covering several complaints gets one structured issue layer each, so a second problem cannot blur into the first.

Statelessness is why that works. The Messages API keeps no memory of its own. Conversational coherence exists only because you pass the complete conversation history in each subsequent request.54 Compaction and trimming are deliberate edits to that resent payload, so whatever you drop is genuinely gone. Two things live outside the exam. The API now offers server-side compaction and context editing.55 And every edit you make near the front of the payload invalidates the prompt cache from that point on.

Position is a design choice. Models process the beginning and the end of a long input reliably and may simply omit findings sitting in middle sections — the "lost in the middle" effect.16 Layout is the mitigation, not exhortation. Assemble each request in layers, on purpose:

Layer Position Lossy? Why there
Case facts, or a key-findings summary of many subagent reports First Never summarized Primacy; carries what every later decision depends on
Summarized older turns Middle Yes, by design Fine for tone and context, never authoritative for numbers
Detailed results under explicit section headers Middle No Headers keep each block addressable despite the position
Trimmed tool results With their turns Trimmed on purpose Relevance, not completeness, is what earns a place in context
Last N turns verbatim Last Never Strongest recency; the live state of the conversation

Tool results dominate context. An order lookup may return 40+ fields when only 5 matter for a return decision, and each one is re-sent on every later request.24 Trim at the tool boundary: a PostToolUse hook is the named interception point that transforms a tool result before the model ever processes it (see 1.5). Apply the same discipline agent to agent. When the downstream context budget is small, require upstream agents to emit structured data (key facts, citations, relevance scores) plus metadata (dates, source locations, methodological context) instead of verbose content and reasoning chains.

How the request payload is reassembled each turnShow that the prompt sent on every request is deliberately assembled from distinct layers in a fixed order — pinned case facts first, summarized older turns in the middle where the lost-in-the-middle effect bites, the last N turns verbatim at the end — and that pinned case facts bypass compaction entirely while older turns and tool results are lossy on purpose.Full session transcriptFullsessiontranscriptCase facts extractorCase factsextractorPinned case facts block (first)Pinned casefacts block(first)Compaction stepCompactionstepSummarized older turns (middle)Summarizedolder turns(middle)Last N turns verbatim (last)Last N turnsverbatim(last)Trimmed tool resultsTrimmedtool resultsAssembled request payloadAssembledrequestpayloadamounts, dates, IDsnever summarizedolder turns onlylossy proserecency windowkeep relevant fieldsread at the toplost-in-the-middlezoneattached to turnsstrongest recency
How the request payload is reassembled each turn

Show that the prompt sent on every request is deliberately assembled from distinct layers in a fixed order — pinned case facts first, summarized older turns in the middle where the lost-in-the-middle effect bites, the last N turns verbatim at the end — and that pinned case facts bypass compaction entirely while older turns and tool results are lossy on purpose.

one transition at a time

Click a layer to see whether it is lossless or lossy on the next request.

Naive concatenation vs. structured aggregationMake clear that the cure for the "lost in the middle" effect is input layout — summary first, addressable sections below — rather than a stronger instruction to the model.Naive concatenationNaiveconcatenationSubagent report 1Subagent report 1Subagent report 2 (middle)Subagent report 2(middle)Subagent report 3Subagent report 3Answer omits middle findingsAnswer omitsmiddle findingsStructured layoutStructured layoutKey findings summary firstKey findingssummary firstSectioned detail with headersSectioned detailwith headersAnswer covers all reportsAnswer covers allreportsstart, read reliablyburiedend, read reliablylost in the middleevery report, condensedaddressable blocksprimacy positionretrievable by header
Naive concatenation vs. structured aggregation

Make clear that the cure for the "lost in the middle" effect is input layout — summary first, addressable sections below — rather than a stronger instruction to the model.

A case-facts block that survives compaction

Scenario 1 · Customer Support Resolution Agent

A billing-dispute conversation runs 40 turns and the oldest turns get summarized. The disputed amount, the charge date and the promise the agent already made must not be inside that summary. The support agent extracts them into a structured case-facts object after every tool call and re-injects it, verbatim and first, on each request. The compacted history stays useful for tone and context; the facts block stays authoritative for arithmetic and policy.

typescript
type CaseFacts = {
  orderId: string;
  orderTotal: string;      // "USD 847.32" — never "about 850 dollars"
  chargeDate: string;      // "2026-03-14"
  refundStatus: string;    // "requested | approved | processed"
  agentCommitments: string[]; // "refund within 5 business days"
};

// Rebuilt on every API request: facts pinned, older turns compacted.
function buildMessages(facts: CaseFacts, olderSummary: string, recentTurns: Message[]) {
  const pinned = [
    '## CASE FACTS (authoritative, never summarized)',
    JSON.stringify(facts, null, 2),
    '## EARLIER CONVERSATION (compacted — may be imprecise)',
    olderSummary,
  ].join('\n\n');

  return [{ role: 'user' as const, content: pinned }, ...recentTurns];
}
Facts are pinned; only the older history is compacted. Exam-shaped — note that pinning first also rewrites the cached prefix each turn.

Trimming lookup_order before it enters the transcript

Scenario 1 · Customer Support Resolution Agent

lookup_order returns 40+ fields (warehouse routing, carrier scan events, marketing attribution). For a return decision only a handful are relevant. Trimming happens in a PostToolUse hook — the interception point for transforming a tool result before the model processes it (see 1.5) — so the full payload never reaches the conversation. The alternative is paying for those fields on every subsequent request for the rest of the session, and burying the fields that matter among ones that do not.

typescript
const RETURN_RELEVANT = [
  'order_id',
  'order_date',
  'order_total',
  'currency',
  'status',
  'items',
  'return_window_ends',
] as const;

/** The upstream response has 40+ fields; only these enter context. */
export function trimOrder(raw: Record<string, unknown>) {
  return Object.fromEntries(
    RETURN_RELEVANT.filter((k) => k in raw).map((k) => [k, raw[k]]),
  );
}
Field allow-list applied in a PostToolUse hook.

Laying out an aggregated research brief

Scenario 3 · Multi-Agent Research System

The coordinator receives four subagent reports and must synthesize them. Concatenating them in arrival order buries report 2 and 3 in the middle of a very long input. Instead the coordinator writes a key-findings block first, then the details under explicit headers, and requires each subagent to have supplied dates, source locations and methodological context so the synthesis step does not have to guess.

text
# KEY FINDINGS (read first — one line per subagent)
1. Web search: adoption rose 34% in 2025 (3 sources, most recent 2026-01).
2. Document analysis: internal figures disagree with public ones (see S2).
3. Filings review: no disclosure found for FY2025 (valid empty result).
4. Competitor scan: two of five competitors ship the feature.

# SECTION 1 — Web search findings
  source | date collected | claim | relevance
  ...

# SECTION 2 — Document analysis findings
  document | publication date | excerpt | claim
  ...

# OPEN QUESTIONS AND COVERAGE GAPS
- FY2025 filings unavailable; conclusions about FY2025 are unsupported.
Aggregated-input template that mitigates position effects.
  • Truncating or summarizing the oldest turns blindly instead of first extracting transactional facts and prior commitments into a pinned layer because compaction destroys precisely the numbers, dates and expectations the agent must honor.
  • Concatenating subagent reports in arrival order instead of leading with a key-findings summary and sectioning the detail because middle-of-input findings are the ones the model omits.
  • Piping raw tool responses into the transcript instead of trimming them to the relevant fields because irrelevant fields are re-sent on every later request and crowd out the relevant ones.
  • Having upstream agents forward verbose content and reasoning chains instead of structured facts, citations and metadata because a downstream agent with a small context budget cannot afford narrative it did not ask for.
  • Stems describe a long support or research session where the agent "forgot" an amount, a date or a promise. The tempting-but-wrong option is a bigger context window or a better summarization prompt; the credited answer extracts the facts into a persistent layer outside the summarized history.
  • When a stem mentions that findings from the middle of a long aggregated input were missing from the final answer, the mechanism being tested is "lost in the middle" — pick reordering plus explicit section headers, not a re-read instruction.
  • Watch for token-cost stems built on a tool returning dozens of fields. The credited answer trims at the tool boundary; distractors trim after the fact, or compact the whole conversation, which also loses the case facts.
Beyond the exam — what the API does that this does not grade

Prompt caching is a prefix match: one changed byte invalidates every breakpoint after it. A facts block that mutates each turn, placed first, therefore writes a new cache on every request and reads none. In production you keep the frozen system prompt and tool list at the front behind a cache_control breakpoint. Deliver mutable state after the cached history — as a {"role": "system", ...} message on models that support it, or as text in the latest user turn otherwise.

The API also offers server-side alternatives to hand-rolled compaction. context_management with clear_tool_uses_20250919 clears old tool results, and compaction summarizes history server-side.55 None of this is tested; all of it decides your bill.

References — 4 sources
  1. Effective context engineering for AI agents Anthropic Engineering, 29 Sep 2025 Names and explains context rot — recall degrades as the window fills — which is the evidence that a bigger window buys capacity, not attention.
  2. Manage tool context Anthropic The four documented remedies for exactly the pressure the 4–5 versus 18 rule describes: tool search, programmatic tool calling, prompt caching, and context editing. It answers the question the rule raises and closes off — what to do when you genuinely need 18 tools.
  3. Context windows Anthropic The normative statement of API statelessness and how the window accumulates across turns, which is the fact every trimming decision in 5.1 depends on.
  4. Context editing Anthropic The API-native alternative to hand-rolled trimming — server-side tool-result clearing with `clear_tool_uses_20250919` — for readers on the Messages API, where a `PostToolUse` hook does not exist.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • Progressive summarization risks: condensing numerical values, percentages, dates, and customer- stated expectations into vague summaries
  • The "lost in the middle" effect: models reliably process information at the beginning and end of long inputs but may omit findings from middle sections
  • How tool results accumulate in context and consume tokens disproportionately to their relevance (e.g., 40+ fields per order lookup when only 5 are relevant)
  • The importance of passing complete conversation history in subsequent API requests to maintain conversational coherence

Skills in

  • Extracting transactional facts (amounts, dates, order numbers, statuses) into a persistent "case facts" block included in each prompt, outside summarized history
  • Extracting and persisting structured issue data (order IDs, amounts, statuses) into a separate context layer for multi-issue sessions
  • Trimming verbose tool outputs to only relevant fields before they accumulate in context (e.g., keeping only return-relevant fields from order lookups)
  • Placing key findings summaries at the beginning of aggregated inputs and organizing detailed results with explicit section headers to mitigate position effects
  • Requiring subagents to include metadata (dates, source locations, methodological context) in structured outputs to support accurate downstream synthesis
  • Modifying upstream agents to return structured data (key facts, citations, relevance scores) instead of verbose content and reasoning chains when downstream agents have limited context budgets
Back to top