Note on selection
Manage context effectively in large codebase exploration
What you need to know
- The tell-tale sign of context degradation is answers drifting from specific discovered classes to generic "typical patterns", plus inconsistency between turns.
- Scratchpad files persist key findings across context boundaries; the transcript does not.
- Delegate bounded investigations to subagents to isolate verbose exploration output and keep the main agent on coordination.
- Summarize each exploration phase and inject that summary into the next phase initial context instead of letting the next agents rediscover.
- Use /compact when the window fills with discovery output — after key findings are written down.
- Crash recovery = each agent exports state to a known location and the coordinator loads a manifest on resume.
Exploring an unfamiliar codebase is the workload that exhausts context fastest, because discovery output is voluminous and most of it is disposable.
Recognize context degradation. The symptom in extended sessions is specific and testable. Answers turn inconsistent across turns, and the model starts referencing "typical patterns" — "usually a repository layer handles this" — instead of the concrete classes and files it discovered earlier.63 That shift from specific to generic is the signal that key findings have been pushed out of effective context, and no amount of re-asking will bring them back. The remedies are all about moving findings out of the transcript and into durable form.
Scratchpad files. Have the agent maintain a scratchpad file recording key findings — entry points, file paths, class names, invariants — and reference that file for subsequent questions. A finding written to disk survives a context boundary; the same finding held only in the transcript does not.16 This is what lets a later question be answered from the recorded specifics instead of from generic pattern-matching.
Subagent delegation as context isolation. Spawn subagents for bounded investigations ("find all test files", "trace the refund flow dependencies"). The subagent absorbs the verbose exploration output — dozens of file reads and grep dumps — and returns a compact answer, while the main agent's context stays reserved for high-level coordination. The point is isolation of verbosity, not parallel speed.16 At monorepo scale the configuration side of the same problem matters too: nested CLAUDE.md files, sparse worktrees and per-package skills are what keep a session from needing these rescue tactics at all.32
Phase boundaries and compaction. Before spawning the next phase's subagents, summarize the key findings of the phase that just finished. Inject that summary into the new agents' initial context so they start informed rather than rediscovering. Within an interactive Claude Code session, use /compact to reduce context usage once the window has filled with verbose discovery output — after the findings you care about are safely in the scratchpad, not before.64
Crash recovery via manifests. Long unattended runs need structured state persistence. Each agent exports its state to a known location; on resume the coordinator loads a manifest and injects the recovered state into agent prompts. That converts a crash from "start the whole exploration again" into "resume the phases that did not complete".
Show the funnel from a broad question to a specific answer, and where verbose output is deliberately kept out of the main agent context.
Click a stage to highlight whether its output lands in a subagent context or the main one.
Show that resumability comes from agents exporting state to known locations and the coordinator loading a manifest and injecting recovered state, rather than from re-running the exploration.
Worked examples
Progressive narrowing instead of reading the repository
Scenario 2 · Code Generation with Claude CodeThe question is "how does the refund flow work?" in a 400k-line service. Reading broadly to build understanding fills the window with code that turns out to be irrelevant and leaves no room for the files that matter. Use the built-in tools in order of cost: Glob for path patterns and Grep for content search (neither loads a file body), then delegated investigations, then a handful of targeted Read calls guided by what the scratchpad says.
# 1. Breadth — path patterns first, then content, no file bodies loaded
Glob(pattern="src/**/*efund*.ts") # path matching: Refund and refund files
Grep(pattern="class .*Refund|function refund",
glob="src/**/*.ts",
output_mode="files_with_matches") # content search: where it lives
# 2. Delegate the verbose parts to subagents (their context, not yours):
# - "trace refund flow dependencies from RefundService to the gateway"
# - "find all test files covering refunds and list what they assert"
# 3. Record what came back in notes/refund-flow.md, then work from the file
- Entry: src/api/refunds/handler.ts -> RefundService.create()
- Policy gate: src/domain/refund/policy.ts (RefundWindowPolicy, 30 days)
- Gateway: src/infra/payments/StripeRefundClient.ts (idempotency key required)
- Tests: test/refund/*.spec.ts (12 files), no test for partial refunds
# 4. Only now Read the 3 files that actually matter, then /compactPhase summary injected into the next phase
Scenario 2 · Code Generation with Claude CodePhase 1 mapped the modules; phase 2 must assess test coverage. Rather than continuing in a context already full of phase-1 grep output, the coordinator condenses phase 1 into a short findings block and passes it as the opening context of each phase-2 subagent. The subagents inherit the specifics — real class names, real paths — which is exactly what degraded context loses.
PHASE 1 FINDINGS (authoritative; do not rediscover)
Refund entry point : src/api/refunds/handler.ts
Core service : src/domain/refund/RefundService.ts
Policy gate : RefundWindowPolicy (30-day window, src/domain/refund/policy.ts)
Payment gateway : StripeRefundClient, requires idempotency key
Known gap : no test covers partial refunds
YOUR TASK (phase 2)
Assess test coverage for the components listed above ONLY.
Write findings to notes/phase2-coverage.md before returning.
Return at most 20 lines: file, covered behavior, gap.Manifest-based crash recovery for a long run
Scenario 3 · Multi-Agent Research SystemAn overnight multi-agent exploration crashes after the third of six phases. Because each agent exported its state to a known path and registered it in a manifest, the coordinator resumes by loading the manifest, injecting completed state into prompts, and re-dispatching only the incomplete phases.
{
"run_id": "explore-2026-08-04T22:10Z",
"updated_at": "2026-08-05T01:47Z",
"phases": [
{ "id": "p1-module-map", "status": "complete",
"state_path": "state/p1-module-map.json", "findings": "notes/modules.md" },
{ "id": "p2-refund-flow", "status": "complete",
"state_path": "state/p2-refund-flow.json", "findings": "notes/refund-flow.md" },
{ "id": "p3-test-coverage", "status": "partial",
"state_path": "state/p3-test-coverage.json",
"completed_units": 8, "total_units": 12,
"resume_hint": "re-dispatch units 9-12 only" },
{ "id": "p4-dependency-risk", "status": "pending", "state_path": null }
],
"inject_on_resume": ["notes/modules.md", "notes/refund-flow.md"]
}Anti-patterns
- Exploring a large codebase by reading files broadly instead of narrowing progressively with search and delegated investigations because the window fills with code that turns out to be irrelevant and the relevant files never fit.
- Keeping key findings only in the conversation instead of writing them to a scratchpad file because context degradation replaces the specific classes discovered earlier with generic "typical pattern" answers.
- Running verbose discovery in the main agent instead of delegating it to subagents because the coordinator loses the high-level context it needs precisely when the exploration gets long.
- Compacting or starting a fresh session to fix degradation without first persisting findings and injecting a phase summary because the next phase then rediscovers everything from scratch.
How it is examined
- The stem often describes the symptom rather than the cause: inconsistent answers and references to "typical patterns" late in a session. Name it as context degradation and pick scratchpad persistence plus delegation.
- Distinguish the two roles of subagents in this domain: here they exist to isolate verbose output, not to parallelize. Options praising speed are usually the weaker choice.
- Crash-recovery items reward structured state exports plus a coordinator-loaded manifest; distractors offer longer timeouts, bigger context windows, or simply restarting the run.
References — 4 sources
- Effective context engineering for AI agents Anthropic Engineering, 29 Sep 2025 Names and explains context rot — recall degrades as the window fills — which is the evidence that a bigger window buys capacity, not attention.
- Set up Claude Code in a monorepo or large codebase Anthropic What changes at monorepo scale — code intelligence, sparse worktrees, per-package configuration — which is where the two-pass grep stops being sufficient.
- Explore the context window Anthropic An interactive simulation of what fills the window during a real session — what loads automatically, what each file read costs, what survives compaction.
- Commands Anthropic The reference entry for `/compact` and its focus instructions, plus the adjacent commands: `/context` to see what is filling the window, `/clear` when compaction is the wrong tool.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- Context degradation in extended sessions: models start giving inconsistent answers and referencing "typical patterns" rather than specific classes discovered earlier
- The role of scratchpad files for persisting key findings across context boundaries
- Subagent delegation for isolating verbose exploration output while the main agent coordinates high- level understanding
- Structured state persistence for crash recovery: each agent exports state to a known location, and the coordinator loads a manifest on resume
Skills in
- Spawning subagents to investigate specific questions (e.g., "find all test files," "trace refund flow dependencies") while the main agent preserves high-level coordination
- Having agents maintain scratchpad files recording key findings, referencing them for subsequent questions to counteract context degradation
- Summarizing key findings from one exploration phase before spawning sub-agents for the next phase, injecting summaries into initial context
- Designing crash recovery using structured agent state exports (manifests) that the coordinator loads on resume and injects into agent prompts
- Using /compact to reduce context usage during extended exploration sessions when context fills with verbose discovery output