Note on selection
Design efficient batch processing strategies
What you need to know
- Batch is a cost-for-latency trade: 50% cheaper, up to 24 hours, no latency guarantee.
- Batch fits non-blocking latency-tolerant work (overnight reports, weekly audits, nightly test generation); blocking pre-merge checks stay on the synchronous API.
- The guide's line is that batch does not support multi-turn tool calling within a single request — that is the credited answer, and the design point behind it is real: batch is for work you submit and walk away from, not for an interactive loop. What actually differs at the API is streaming, which batch does not support; tool use itself, including server tools, is available in a batch request.
- custom_id is the correlation key for request/response pairs and the handle you use to resubmit only what failed.
- Results are retrieved by polling for completion — there is no push callback, and polling makes nothing arrive sooner, so it never makes batch safe for a blocking gate.
- The guide's cadence arithmetic = promised SLA minus the 24-hour window; a 30-hour SLA implies submitting at least every 4 hours. The window is an expiry, not a delivery promise: unfinished requests come back expired and need another cycle.
- Refine the prompt on a sample set before batching a large corpus, so first-pass success is high and resubmission costs stay low.
What the Message Batches API buys and what it costs
The trade is explicit, and it runs on more axes than cost.
| Synchronous Messages API | Message Batches API | |
|---|---|---|
| Cost | standard pricing | 50% cost savings |
| Latency | immediate response | up to 24 hours, no guaranteed SLA |
| Tool calling | client-side tool loops run turn by turn, and the response streams | same turn-by-turn loop, but no stream parameter — you retrieve a finished result, not a live one |
| Retrieval | the response is the return value | poll for completion; no push callback |
| Workload it fits | blocking work: pre-merge gates, interactive commands | non-blocking, latency-tolerant work: overnight quality reports, weekly compliance audits, nightly test generation, corpus backfills |
Most batches finish far sooner than the window, but you cannot build a promise on that. The only number you may plan against is 24 hours, which is why a pre-merge check that gates a developer's merge cannot sit in that queue. One operational detail beyond the guide: at the 24-hour mark, unfinished requests come back marked expired, not late.50 The ceiling is a deadline, not a delay.
The capability limit deserves reading precisely, because the exam tests its exact scope. The guide excludes the multi-turn case, so an agentic loop that needs to call a tool, read the result and continue belongs on the synchronous API. The wording implies that a one-shot request returning a tool_use block — the structured-extraction pattern from 4.3 — is not what the restriction is about, since no tool is executed mid-request. The API documentation settles that reading: tool use, including all server tools, and multi-turn conversations are both batchable.50 What is testable is the guide's exclusion itself, so answer with it on the exam.
custom_id is the join key, and results are polled
Each request carries a custom_id and every result comes back with it. Results are not guaranteed to arrive in submission order, so custom_id — not array position — is how you correlate response to request, and how you identify precisely which documents failed.51 Keep the id short and plain: the API constrains its characters and its length.52 Each result also carries a type. succeeded and expired are the two that matter for the certification; the API also returns others specific to it.
Retrieval is a pull, not a push: you check the batch's status until it is done, then read the results. Nothing wakes your pipeline up when the work finishes, which is a second reason batch cannot sit under a blocking gate. Adding status polling does not rescue such a workflow either, because progress reporting shortens nothing.
Sizing submission cadence against an SLA
The guide's own worked example: you promise a 30-hour SLA and batch processing may take up to 24 hours. That leaves 6 hours of slack, which is the maximum time a document may wait in your queue before submission. Submitting every 4 hours therefore keeps the worst case at 4 + 24 = 28 hours, inside the promise with margin. That is the guide's arithmetic and the credited answer; it assumes the window delivers, which is what the expiry note above corrects. The general recipe: maximum queue window = promised SLA − 24 hours, then pick a submission interval at or below that.
Handling failures and de-risking the first pass
When a batch ends with partial failures, resubmit only the failed custom_ids, with a modification appropriate to the cause — chunk the documents that exceeded context limits, fix the ones that hit a validation problem. Resubmitting the whole batch pays a second time for work that already succeeded.
Because a wrong prompt is amplified by volume, refine the prompt on a small sample set before batch-processing large volumes. Every first-pass failure becomes an iterative resubmission cost, so a sample run that lifts the first-pass success rate is usually the cheapest step in the whole pipeline.
Give the learner a two-row table of the two APIs against the four columns the exam tests — cost, latency guarantee, multi-turn tool calling and workload fit — so a stem can be mapped to an API in one lookup.
Make the SLA arithmetic visual: the queue wait plus the 24-hour batch window must fit inside the promised SLA, which is what fixes the submission interval.
Worked examples
Splitting a CI workload between the two APIs
Scenario 5 · Claude Code for Continuous IntegrationThe instinct to move everything to batches for the 50% discount is the trap. The pre-merge review blocks a human and must be synchronous; the nightly test generation and the weekly audit block nobody and are exactly what batches are for. The answer is a split, not a single API.
Pre-merge PR review blocking, developer waiting -> synchronous API
Post-merge deep analysis non-blocking, results by 9am -> Message Batches API
Nightly test generation non-blocking, overnight -> Message Batches API
Weekly compliance audit non-blocking, 7-day cadence -> Message Batches API
Interactive /explain cmd blocking, human in the loop -> synchronous API
Rule of thumb: if a person or a pipeline gate is waiting on the answer,
there is no latency SLA to rely on and batch is the wrong choice --
regardless of the 50% saving.Selective resubmission keyed on custom_id
Scenario 6 · Structured Data ExtractionResults are consumed by custom_id, never by position. Failures are collected and resubmitted individually with a cause-appropriate modification — here, documents that exceeded the context limit are chunked before the retry, while other failures are resubmitted unchanged. Nothing that already succeeded is paid for twice.
const failed: string[] = [];
for await (const r of client.messages.batches.results(batchId)) {
if (r.result.type === 'succeeded') {
save(r.custom_id, r.result.message); // key by custom_id, not order
} else {
failed.push(r.custom_id); // errored | expired | canceled
}
}
// Resubmit only the failures, with a modification matched to the cause.
const retryRequests = failed.flatMap((id) =>
exceededContextLimit(id)
? chunk(documents[id]).map((part, i) => buildRequest(id + '-part' + i, part))
: [buildRequest(id, documents[id])]
);
if (retryRequests.length > 0) {
await client.messages.batches.create({ requests: retryRequests });
}Deriving submission frequency from a promised SLA
Scenario 6 · Structured Data ExtractionThe arithmetic the exam expects. Only the 24-hour worst case is contractual, so the entire remaining budget is queue time, and the submission interval must fit inside it.
Promised end-to-end SLA .................... 30 h
Batch processing worst case ................ 24 h
Remaining budget for queue wait ............ 6 h
Submission interval chosen .................. 4 h
Worst case for any document ..... 4 h wait + 24 h processing = 28 h <= 30 h OK
Submitting once daily (24 h interval) gives 24 + 24 = 48 h -- breaks the SLA
even though most batches would finish in under an hour. Plan against the
documented window, not against observed latency.
The exam's arithmetic ends here. In production add margin: a request still
running at 24 h is returned as `expired` and must be resubmitted, so the real
worst case is 24 h + one more submission cycle.Anti-patterns
- Batching a blocking gate: routing pre-merge PR review through the Message Batches API to capture the 50% discount because there is no latency SLA and the merge could wait up to 24 hours.
- Full-corpus resubmission: resubmitting the entire batch after partial failures instead of only the failed custom_ids with a cause-appropriate fix because you pay a second time for work that already succeeded.
- Batching an agentic loop: putting a multi-turn tool-calling workflow into a single batch request because the guide excludes that case and batch is for work you submit and walk away from.
- Batching before refining: submitting a large corpus before validating the prompt on a sample set because every first-pass failure becomes an iterative resubmission cost that dwarfs the sample run.
How it is examined
- Stems give you a latency requirement and a cost pressure at the same time. The credited answer usually splits the workload — synchronous for the blocking path, batch for the overnight or weekly path — rather than choosing one API for everything.
- Expect SLA arithmetic: subtract the 24-hour batch window from the promised SLA and the remainder is your maximum submission interval. Answers computed from typical observed latency are wrong.
- An option that runs a multi-turn tool-calling loop inside a batch request is wrong on the guide's exclusion alone, even if its latency profile looks acceptable. Do not argue from the API, where tool use is batchable and only streaming is missing.
Beyond the exam — what the API does that this does not grade
Yes for the requests, no for the interactivity. Tool use is batchable, including all server tools, and so are multi-turn conversations.50 What batch drops is stream, plus a short list of other parameters. You retrieve a finished result instead of watching one arrive.
A client-side tool loop is many requests on either API: call, run the tool locally, call again. Batch does not change that shape. It changes when each answer comes back. So submit the turns that do not depend on each other, then correlate them by custom_id. Anything a person or a gate is waiting on stays on the synchronous API.
References — 3 sources
- Batch processing Anthropic 24 hours is an expiry, not a completion time: "Batches expire if processing does not complete within 24 hours", and an expired request returns no result and is not billed. The same page settles the tool question — "Tool use, including all server tools" and "Multi-turn conversations" are listed under What can be batched, and the real exclusion list is `stream`, `speed`, `store`, `previous_thread_event_id`, `cache_hint`, `context_hint`, `max_tokens: 0`.
- Retrieve Message Batch results Anthropic The API reference in normative form — "Batch results can be returned in any order… always use the `custom_id` field" — plus the `succeeded` / `errored` / `canceled` / `expired` result types this unit's resubmission logic has to branch on.
- Create a Message Batch Anthropic The actual `custom_id` constraint — 1–64 characters matching `^[a-zA-Z0-9_-]{1,64}$` — which rules out the document paths and URLs a reader would naturally reach for as join keys.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- The Message Batches API: 50% cost savings, up to 24-hour processing window, no guaranteed latency SLA
- Batch processing is appropriate for non-blocking, latency-tolerant workloads (overnight reports, weekly audits, nightly test generation) and inappropriate for blocking workflows (pre-merge checks)
- The batch API does not support multi-turn tool calling within a single request (cannot execute tools mid-request and return results)
- custom_id fields for correlating batch request/response pairs
Skills in
- Matching API approach to workflow latency requirements: synchronous API for blocking pre-merge checks, batch API for overnight/weekly analysis
- Calculating batch submission frequency based on SLA constraints (e.g., 4-hour windows to guarantee 30-hour SLA with 24-hour batch processing)
- Handling batch failures: resubmitting only failed documents (identified by custom_id) with appropriate modifications (e.g., chunking documents that exceeded context limits)
- Using prompt refinement on a sample set before batch-processing large volumes to maximize first- pass success rates and reduce iterative resubmission costs