Note on selection
Preserve information provenance and handle uncertainty in multi- source synthesis
What you need to know
- Claim-source mappings (source URL or document name plus relevant excerpt) must be created at discovery and preserved and merged through synthesis, never reconstructed later.
- Conflicting statistics from credible sources are annotated with attribution and passed up for the coordinator to reconcile — never resolved by arbitrary selection.
- Reports separate well-established from contested findings and keep each source original characterization and methodological context.
- Publication and data-collection dates in structured output stop temporal differences from being misread as contradictions.
- Match rendering to content type — forcing tables, prose and findings into one shape loses the structure.
When many sources are merged into one deliverable, two things are routinely destroyed: where each claim came from, and how certain it is. Both losses happen at the same place — the summarization step.
Attribution dies in compression, and the loss is one-way. A subagent finds a figure in a named report and writes "adoption grew 34%". Once that sentence is summarized into a coordinator-level brief, the URL, document name and excerpt are gone. No downstream agent can restore them, so the synthesis agent either drops the citation or invents a plausible one. The API-native answer is model-generated citations bound to spans of a supplied document, returned as structured blocks.66 This is why provenance has to travel as data through synthesis rather than be reconstructed from prose afterwards: every downstream agent preserves and merges the structured record instead of rewriting it. Anthropic's own multi-agent research system had to solve this at production scale, and runs a separate citation pass to do it.4
Conflicts are annotated, not resolved by fiat. When two credible sources give different statistics, silently choosing one erases a real disagreement and makes the report look more settled than the evidence is. The general principle is to give the model an explicit way to express uncertainty rather than force a confident answer.49 Complete the analysis with both values included and attributed, and let the coordinator reconcile before anything reaches synthesis. The report then separates well-established findings from contested ones, which is the certainty gradient a reader needs in order to weigh a claim.
What every finding carries, and what breaks without it. Each element is a separate failure mode, so a record missing one is not merely less detailed:
| Element in the structured finding | What it preserves | What breaks without it |
|---|---|---|
| Source URL or document name, plus the relevant excerpt | The claim-source mapping, still verifiable after synthesis | Citations become guesses — dropped, or invented to look plausible |
| Publication and data-collection dates | Which vintage of a measurement this is | A 2023 figure and a 2026 figure read as a contradiction instead of a time series |
| Methodology and each source's original characterization | Why two credible numbers can both be right | "Self-reported survey of 400 firms" and "audited filings" collapse into one unexplained gap |
| Both values when sources conflict, each attributed | A real disagreement, visible as a disagreement | Arbitrary selection overstates certainty and hides the conflict from the coordinator |
| Content type — financial data as tables, news as prose, technical findings as structured lists | Comparability of numbers, nuance of qualitative material | One uniform format loses both: narrative paragraphs bury the numbers, bullet fragments shred the nuance |
Contrast the path that carries claim-source mappings as structured data through synthesis with the path that compresses findings into prose and loses attribution irreversibly.
Show why keeping both conflicting values with attribution, methodology and dates is the only option that lets the coordinator reconcile knowingly.
Worked examples
A claim-source mapping that survives synthesis
Scenario 3 · Multi-Agent Research SystemEach research subagent returns findings as records, not prose. The coordinator merges the arrays; the synthesis agent is instructed to carry the source, excerpt and date fields through into the report and to refuse to state any claim that has lost its mapping. That constraint is what makes the final citations verifiable rather than decorative.
{
"findings": [
{
"claim": "EU adoption of the technology grew 34% year over year",
"source": "https://example-institute.org/reports/eu-adoption-2026",
"source_name": "EU Adoption Monitor 2026",
"excerpt": "adoption across the EU-27 rose 34% relative to 2025",
"publication_date": "2026-02-11",
"data_collection_period": "2025-07 to 2025-12",
"methodology": "self-reported survey, n=412 firms",
"relevance_score": 0.91
}
],
"synthesis_contract": "preserve claim, source, excerpt, dates and methodology verbatim; do not merge claims from different sources into one sentence"
}Two credible sources, two different numbers
Scenario 3 · Multi-Agent Research SystemThe document analyst finds 34% in a survey report and 21% in audited filings. Picking either one would hide a real methodological difference. The analyst completes its task with both values annotated, and the coordinator — which can see the methodology and the collection periods — decides whether they are a conflict, a time-series difference, or two measurements of different populations.
{
"metric": "year-over-year adoption growth, EU",
"status": "conflicting_values",
"values": [
{ "value": "34%", "source_name": "EU Adoption Monitor 2026",
"publication_date": "2026-02-11", "collection_period": "2025-H2",
"methodology": "self-reported survey, n=412" },
{ "value": "21%", "source_name": "Regulator Filing Digest",
"publication_date": "2026-05-30", "collection_period": "FY2025",
"methodology": "audited filings, n=1,180" }
],
"analyst_note": "different populations and periods; likely not a contradiction",
"resolution": "deferred to coordinator; do not average or drop either value"
}Report structure that separates established from contested
Scenario 6 · Structured Data ExtractionThe extraction and reporting layer renders each content type in its natural form and keeps the certainty gradient visible. Financial figures go in a table where values are comparable; contested claims get their own section with both characterizations preserved; qualitative context stays as prose instead of being shredded into bullets.
## Well-established findings
| Metric | Value | Source | Collected | Method |
|---|---|---|---|---|
| Revenue FY2025 | USD 1.28B | Annual Report 2025 | FY2025 | audited |
| Headcount | 4,120 | Annual Report 2025 | 2025-12-31 | audited |
## Contested findings
**YoY adoption growth — 34% vs 21%.** The EU Adoption Monitor 2026
(self-reported survey, n=412, collected 2025-H2) reports 34%. The Regulator
Filing Digest (audited filings, n=1,180, FY2025) reports 21%. The populations
and periods differ; both figures are retained.
## Context and developments (prose)
Coverage in Q1 2026 focused on the compliance deadline; no source disputes
the timeline itself, only the cost estimates attached to it.Anti-patterns
- Summarizing subagent findings into prose instead of preserving structured claim-source mappings because once the URL, document name and excerpt are compressed away no downstream agent can restore them and citations become guesses.
- Arbitrarily selecting one of two conflicting statistics instead of annotating both with source attribution and deferring reconciliation to the coordinator because it erases a real disagreement and overstates certainty.
- Omitting publication or collection dates from structured outputs because differently dated measurements of the same metric then look like sources contradicting each other.
- Converting every source into one uniform format instead of rendering financial data as tables, news as prose and technical findings as structured lists because flattening destroys comparability and methodological nuance.
How it is examined
- When a stem says the final report cited sources incorrectly or generically, the mechanism is attribution lost during an intermediate summarization step — the credited answer adds structured claim-source mappings that downstream agents must preserve.
- For conflicting-statistics stems, every option that picks a winner (most recent, most reputable, the average) is wrong: annotate both with attribution and let the coordinator reconcile.
- If the stem mentions figures that "contradict" each other and the sources are from different years, the answer is requiring publication or data-collection dates, not adjudicating the numbers.
References — 3 sources
- How we built our multi-agent research system Anthropic Engineering, 13 Jun 2025 The field report where narrow decomposition was observed and measured — one subagent on the 2021 chip crisis while two duplicated work on 2025 supply chains.
- Reduce hallucinations Anthropic Anthropic's guidance on giving the model a legal way to say "not present", which is the general principle the nullable-field rule implements at the schema layer.
- Citations Anthropic Model-generated citations bound to spans of a supplied document, returned as structured blocks — the artifact that makes an ambiguity flag checkable rather than a judgment call.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- How source attribution is lost during summarization steps when findings are compressed without preserving claim-source mappings
- The importance of structured claim-source mappings that the synthesis agent must preserve and merge when combining findings
- How to handle conflicting statistics from credible sources: annotating conflicts with source attribution rather than arbitrarily selecting one value
- Temporal data: requiring publication/collection dates in structured outputs to prevent temporal differences from being misinterpreted as contradictions
Skills in
- Requiring subagents to output structured claim-source mappings (source URLs, document names, relevant excerpts) that downstream agents preserve through synthesis
- Structuring reports with explicit sections distinguishing well-established findings from contested ones, preserving original source characterizations and methodological context
- Completing document analysis with conflicting values included and explicitly annotated, letting the coordinator decide how to reconcile before passing to synthesis
- Requiring subagents to include publication or data collection dates in structured outputs to enable correct temporal interpretation
- Rendering different content types appropriately in synthesis outputs—financial data as tables, news as prose, technical findings as structured lists—rather than converting everything to a uniform format