Skip to content
CCAR-FAcademy
Domain 4 · Statement 4.5 5 of 6
4.5

Design efficient batch processing strategies

  • Batch is a cost-for-latency trade: 50% cheaper, up to 24 hours, no latency guarantee.
  • Batch fits non-blocking latency-tolerant work (overnight reports, weekly audits, nightly test generation); blocking pre-merge checks stay on the synchronous API.
  • The guide's line is that batch does not support multi-turn tool calling within a single request — that is the credited answer, and the design point behind it is real: batch is for work you submit and walk away from, not for an interactive loop. What actually differs at the API is streaming, which batch does not support; tool use itself, including server tools, is available in a batch request.
  • custom_id is the correlation key for request/response pairs and the handle you use to resubmit only what failed.
  • Results are retrieved by polling for completion — there is no push callback, and polling makes nothing arrive sooner, so it never makes batch safe for a blocking gate.
  • The guide's cadence arithmetic = promised SLA minus the 24-hour window; a 30-hour SLA implies submitting at least every 4 hours. The window is an expiry, not a delivery promise: unfinished requests come back expired and need another cycle.
  • Refine the prompt on a sample set before batching a large corpus, so first-pass success is high and resubmission costs stay low.

What the Message Batches API buys and what it costs

The trade is explicit, and it runs on more axes than cost.

Synchronous Messages API Message Batches API
Cost standard pricing 50% cost savings
Latency immediate response up to 24 hours, no guaranteed SLA
Tool calling client-side tool loops run turn by turn, and the response streams same turn-by-turn loop, but no stream parameter — you retrieve a finished result, not a live one
Retrieval the response is the return value poll for completion; no push callback
Workload it fits blocking work: pre-merge gates, interactive commands non-blocking, latency-tolerant work: overnight quality reports, weekly compliance audits, nightly test generation, corpus backfills

Most batches finish far sooner than the window, but you cannot build a promise on that. The only number you may plan against is 24 hours, which is why a pre-merge check that gates a developer's merge cannot sit in that queue. One operational detail beyond the guide: at the 24-hour mark, unfinished requests come back marked expired, not late.50 The ceiling is a deadline, not a delay.

The capability limit deserves reading precisely, because the exam tests its exact scope. The guide excludes the multi-turn case, so an agentic loop that needs to call a tool, read the result and continue belongs on the synchronous API. The wording implies that a one-shot request returning a tool_use block — the structured-extraction pattern from 4.3 — is not what the restriction is about, since no tool is executed mid-request. The API documentation settles that reading: tool use, including all server tools, and multi-turn conversations are both batchable.50 What is testable is the guide's exclusion itself, so answer with it on the exam.

custom_id is the join key, and results are polled

Each request carries a custom_id and every result comes back with it. Results are not guaranteed to arrive in submission order, so custom_id — not array position — is how you correlate response to request, and how you identify precisely which documents failed.51 Keep the id short and plain: the API constrains its characters and its length.52 Each result also carries a type. succeeded and expired are the two that matter for the certification; the API also returns others specific to it.

Retrieval is a pull, not a push: you check the batch's status until it is done, then read the results. Nothing wakes your pipeline up when the work finishes, which is a second reason batch cannot sit under a blocking gate. Adding status polling does not rescue such a workflow either, because progress reporting shortens nothing.

Sizing submission cadence against an SLA

The guide's own worked example: you promise a 30-hour SLA and batch processing may take up to 24 hours. That leaves 6 hours of slack, which is the maximum time a document may wait in your queue before submission. Submitting every 4 hours therefore keeps the worst case at 4 + 24 = 28 hours, inside the promise with margin. That is the guide's arithmetic and the credited answer; it assumes the window delivers, which is what the expiry note above corrects. The general recipe: maximum queue window = promised SLA − 24 hours, then pick a submission interval at or below that.

Handling failures and de-risking the first pass

When a batch ends with partial failures, resubmit only the failed custom_ids, with a modification appropriate to the cause — chunk the documents that exceeded context limits, fix the ones that hit a validation problem. Resubmitting the whole batch pays a second time for work that already succeeded.

Because a wrong prompt is amplified by volume, refine the prompt on a small sample set before batch-processing large volumes. Every first-pass failure becomes an iterative resubmission cost, so a sample run that lifts the first-pass success rate is usually the cheapest step in the whole pipeline.

Synchronous API vs Message Batches APIGive the learner a two-row table of the two APIs against the four columns the exam tests — cost, latency guarantee, multi-turn tool calling and workload fit — so a stem can be mapped to an API in one lookup.costlatencycapabilityfitSynchronous Messages APISynchronous MessagesAPIMessage Batches APIMessage Batches APIStandard pricingStandard pricingImmediate responseImmediate responseMulti-turn tools within one requestMulti-turn tools within onerequestBlocking pre-merge checksBlocking pre-mergechecks50% cost savings50% cost savings24-hour window, no SLA24-hour window, no SLANo multi-turn tool loop per the guide's exclusionNo multi-turn tool loop perthe guide's exclusionOvernight reports and weekly auditsOvernight reports andweekly audits
Synchronous API vs Message Batches API

Give the learner a two-row table of the two APIs against the four columns the exam tests — cost, latency guarantee, multi-turn tool calling and workload fit — so a stem can be mapped to an API in one lookup.

Fitting a 4-hour submission cadence inside a 30-hour SLAMake the SLA arithmetic visual: the queue wait plus the 24-hour batch window must fit inside the promised SLA, which is what fixes the submission interval.Hour 0 document arrivesdocument arrivesHour 4 batch submittedbatch submittedHour 28 window closes, results or expiredwindow closes,results or expiredHour 30 SLA deadlineSLA deadlineQueued up to 4 hourswaits for next submission window ·submission intervalProcessing up to 24 hoursdocumented window · end of the windowtwo hours of margin0h4h28h30h
Fitting a 4-hour submission cadence inside a 30-hour SLA

Make the SLA arithmetic visual: the queue wait plus the 24-hour batch window must fit inside the promised SLA, which is what fixes the submission interval.

Splitting a CI workload between the two APIs

Scenario 5 · Claude Code for Continuous Integration

The instinct to move everything to batches for the 50% discount is the trap. The pre-merge review blocks a human and must be synchronous; the nightly test generation and the weekly audit block nobody and are exactly what batches are for. The answer is a split, not a single API.

text
Pre-merge PR review        blocking, developer waiting   -> synchronous API
Post-merge deep analysis   non-blocking, results by 9am  -> Message Batches API
Nightly test generation    non-blocking, overnight       -> Message Batches API
Weekly compliance audit    non-blocking, 7-day cadence   -> Message Batches API
Interactive /explain cmd   blocking, human in the loop   -> synchronous API

Rule of thumb: if a person or a pipeline gate is waiting on the answer,
there is no latency SLA to rely on and batch is the wrong choice --
regardless of the 50% saving.
Workload-to-API mapping

Selective resubmission keyed on custom_id

Scenario 6 · Structured Data Extraction

Results are consumed by custom_id, never by position. Failures are collected and resubmitted individually with a cause-appropriate modification — here, documents that exceeded the context limit are chunked before the retry, while other failures are resubmitted unchanged. Nothing that already succeeded is paid for twice.

typescript
const failed: string[] = [];

for await (const r of client.messages.batches.results(batchId)) {
  if (r.result.type === 'succeeded') {
    save(r.custom_id, r.result.message);   // key by custom_id, not order
  } else {
    failed.push(r.custom_id);              // errored | expired | canceled
  }
}

// Resubmit only the failures, with a modification matched to the cause.
const retryRequests = failed.flatMap((id) =>
  exceededContextLimit(id)
    ? chunk(documents[id]).map((part, i) => buildRequest(id + '-part' + i, part))
    : [buildRequest(id, documents[id])]
);

if (retryRequests.length > 0) {
  await client.messages.batches.create({ requests: retryRequests });
}
Collect failures, then resubmit only those

Deriving submission frequency from a promised SLA

Scenario 6 · Structured Data Extraction

The arithmetic the exam expects. Only the 24-hour worst case is contractual, so the entire remaining budget is queue time, and the submission interval must fit inside it.

text
Promised end-to-end SLA .................... 30 h
Batch processing worst case ................ 24 h
Remaining budget for queue wait ............  6 h

Submission interval chosen ..................  4 h
Worst case for any document ..... 4 h wait + 24 h processing = 28 h  <= 30 h  OK

Submitting once daily (24 h interval) gives 24 + 24 = 48 h -- breaks the SLA
even though most batches would finish in under an hour. Plan against the
documented window, not against observed latency.

The exam's arithmetic ends here. In production add margin: a request still
running at 24 h is returned as `expired` and must be resubmitted, so the real
worst case is 24 h + one more submission cycle.
SLA arithmetic — the guide's calculation; see the expiry note in the body
  • Batching a blocking gate: routing pre-merge PR review through the Message Batches API to capture the 50% discount because there is no latency SLA and the merge could wait up to 24 hours.
  • Full-corpus resubmission: resubmitting the entire batch after partial failures instead of only the failed custom_ids with a cause-appropriate fix because you pay a second time for work that already succeeded.
  • Batching an agentic loop: putting a multi-turn tool-calling workflow into a single batch request because the guide excludes that case and batch is for work you submit and walk away from.
  • Batching before refining: submitting a large corpus before validating the prompt on a sample set because every first-pass failure becomes an iterative resubmission cost that dwarfs the sample run.
  • Stems give you a latency requirement and a cost pressure at the same time. The credited answer usually splits the workload — synchronous for the blocking path, batch for the overnight or weekly path — rather than choosing one API for everything.
  • Expect SLA arithmetic: subtract the 24-hour batch window from the promised SLA and the remainder is your maximum submission interval. Answers computed from typical observed latency are wrong.
  • An option that runs a multi-turn tool-calling loop inside a batch request is wrong on the guide's exclusion alone, even if its latency profile looks acceptable. Do not argue from the API, where tool use is batchable and only streaming is missing.
Beyond the exam — what the API does that this does not grade

Yes for the requests, no for the interactivity. Tool use is batchable, including all server tools, and so are multi-turn conversations.50 What batch drops is stream, plus a short list of other parameters. You retrieve a finished result instead of watching one arrive.

A client-side tool loop is many requests on either API: call, run the tool locally, call again. Batch does not change that shape. It changes when each answer comes back. So submit the turns that do not depend on each other, then correlate them by custom_id. Anything a person or a gate is waiting on stays on the synchronous API.

References — 3 sources
  1. Batch processing Anthropic 24 hours is an expiry, not a completion time: "Batches expire if processing does not complete within 24 hours", and an expired request returns no result and is not billed. The same page settles the tool question — "Tool use, including all server tools" and "Multi-turn conversations" are listed under What can be batched, and the real exclusion list is `stream`, `speed`, `store`, `previous_thread_event_id`, `cache_hint`, `context_hint`, `max_tokens: 0`.
  2. Retrieve Message Batch results Anthropic The API reference in normative form — "Batch results can be returned in any order… always use the `custom_id` field" — plus the `succeeded` / `errored` / `canceled` / `expired` result types this unit's resubmission logic has to branch on.
  3. Create a Message Batch Anthropic The actual `custom_id` constraint — 1–64 characters matching `^[a-zA-Z0-9_-]{1,64}$` — which rules out the document paths and URLs a reader would naturally reach for as join keys.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • The Message Batches API: 50% cost savings, up to 24-hour processing window, no guaranteed latency SLA
  • Batch processing is appropriate for non-blocking, latency-tolerant workloads (overnight reports, weekly audits, nightly test generation) and inappropriate for blocking workflows (pre-merge checks)
  • The batch API does not support multi-turn tool calling within a single request (cannot execute tools mid-request and return results)
  • custom_id fields for correlating batch request/response pairs

Skills in

  • Matching API approach to workflow latency requirements: synchronous API for blocking pre-merge checks, batch API for overnight/weekly analysis
  • Calculating batch submission frequency based on SLA constraints (e.g., 4-hour windows to guarantee 30-hour SLA with 24-hour batch processing)
  • Handling batch failures: resubmitting only failed documents (identified by custom_id) with appropriate modifications (e.g., chunking documents that exceeded context limits)
  • Using prompt refinement on a sample set before batch-processing large volumes to maximize first- pass success rates and reduce iterative resubmission costs
Back to top