Note on selection
Integrate Claude Code into CI/CD pipelines
What you need to know
- Use
-p(--print) so Claude Code runs non-interactively and the CI job never hangs waiting for input. -
--output-format jsonplus--json-schemaproduce machine-parseable findings a pipeline can post as inline PR comments. - CLAUDE.md is how a CI-invoked run gets testing standards, fixture conventions and review criteria.
- An independent review instance beats the session that wrote the code, which is biased toward its own assumptions.
- Feed prior review findings back in and ask for only new or still-unaddressed issues, so re-runs do not duplicate comments.
- Provide the existing test files so generated tests do not duplicate scenarios already covered.
Running Claude Code in a pipeline changes three things: there is no human to answer prompts, the output must be parsed by a machine, and the session has no accumulated project knowledge unless you give it some.
Surface by surface
| What the pipeline is missing | The mechanism | What it produces |
|---|---|---|
| Nobody is there to answer a prompt | the -p (or --print) flag |
a non-interactive run that takes the prompt, produces output and exits — the flag that prevents a CI job from hanging, and the single most-tested fact in this task statement42 |
| A machine, not a reader, consumes the findings | --output-format json with --json-schema |
structured output under an enforced schema — file, line, severity, message — that a pipeline step loops over and posts as inline pull-request comments at the right locations, instead of one generic paragraph44 |
| The run knows nothing about the project | CLAUDE.md, read straight from the checked-out repository |
testing standards, fixture conventions and review criteria on every invocation |
| Generated tests are trivial, or re-cover ground the suite already holds | the existing test files in context, plus CLAUDE.md on what makes a test valuable and which fixtures already exist |
tests that add coverage and reuse the team's fixtures, rather than getter tests and hand-rolled fixture objects |
| The reviewer helped write what it is reviewing | a separate invocation, never a continuation of the generating session | an independent review instance that sees only the diff and the standards |
| A re-run re-reports everything | the prior review findings in context, with an instruction to report only new or still-unaddressed issues | a PR thread that stays readable across ten pushes |
What matters for the exam is that the two output flags exist and are used together. --output-format is a selector, not a switch: it takes text (the default), json, and stream-json.43 The guide tests json because a CI step needs one parseable document at the end — stream-json emits messages as they arrive, which is for a live consumer, not a posting script. If a stem offers stream-json for a "post findings as PR comments" job, it is the plausible-but-wrong option.
The way a schema is handed to --json-schema in the snippets below — a path to a schema file — is written as an illustration, not as a quoted argument specification.43
This is the CLI/CI layer. At the API layer the equivalent mechanism is tool use with a JSON input_schema (Domain 4, task statement 4.3), which is described there as the most reliable way to guarantee schema-compliant output. Same goal — a shape the caller can rely on — enforced at two different layers, so read the stem for which one it is asking about.
Session context isolation
A subtle but heavily tested point: the same session that generated code is less effective at reviewing its own changes than an independent review instance. The generating session carries the assumptions and justifications that produced the code, so it is primed to consider them correct. A fresh instance judges the diff on its merits, which is why the review step is its own job in the pipeline.45
Noise is a context problem, not a model problem
Two of those rows are the same move made twice: hand the run the context that lets it stay quiet. Undocumented fixtures produce hand-rolled objects and trivial assertions. Missing prior findings produce a thread of repeats. Either way the team learns to ignore the bot, and that is the real cost.
Show the full non-interactive review loop: what context goes in (CLAUDE.md, diff, existing tests, prior findings), which flags shape the output, and how structured findings become inline PR comments without duplicates.
Click a context input to see what noise problem it prevents.
Explain why the session that generated the code is a weaker reviewer than a fresh instance that sees only the diff and the project standards.
Worked examples
Automated PR review step in CI
Scenario 5 · Claude Code for Continuous IntegrationThe review job runs Claude Code with -p so it cannot block on input, and with --output-format json --json-schema so the result is validated structured data. A following step iterates the findings array and posts each one as an inline comment on the exact file and line.
Note what is deliberately not in this job: it is a separate invocation from anything that generated code, so the reviewer has no stake in the implementation. And the review criteria are not inlined in the prompt — they live in CLAUDE.md, where the team maintains them under review.
- name: Claude review
run: |
git diff origin/${{ github.base_ref }}...HEAD > /tmp/diff.patch
claude -p "Review the diff in /tmp/diff.patch against the review criteria \
in CLAUDE.md. Report only actionable defects with file and line." \
--output-format json \
--json-schema .claude/schemas/review-findings.json \
> findings.json
- name: Post inline comments
run: node scripts/post-inline-comments.js findings.jsonA schema that makes findings postable
Scenario 5 · Claude Code for Continuous IntegrationWithout a schema, one run returns { "issues": [...] } and the next returns markdown prose, and the posting script breaks. --json-schema fixes the shape so the pipeline can rely on it, and constraining severity to an enum lets the job fail the build only on high-severity findings — a direct lever on false-positive noise.
{
"type": "object",
"required": ["findings"],
"properties": {
"findings": {
"type": "array",
"items": {
"type": "object",
"required": ["file", "line", "severity", "message"],
"properties": {
"file": { "type": "string" },
"line": { "type": "integer", "minimum": 1 },
"severity": { "enum": ["high", "medium", "low"] },
"message": { "type": "string" },
"isNew": { "type": "boolean" }
},
"additionalProperties": false
}
}
}
}Re-running the review after new commits, and generating tests that add value
Scenario 5 · Claude Code for Continuous IntegrationTwo noise problems, two context fixes.
Duplicate comments. On the second push the job passes the previous findings.json back in and instructs Claude to report only issues that are new or still unaddressed. Resolved findings disappear from the thread instead of being restated.
Low-value tests. The test-generation job supplies the existing test files so Claude can see which scenarios are already covered, and relies on CLAUDE.md for what the team considers a valuable test and which fixtures exist. Without this, the job reliably produces getter tests and hand-rolled fixture objects that duplicate tests/fixtures/.
# Re-review: pass prior findings so nothing is reported twice
claude -p "Prior findings are in prior-findings.json. Review the current diff \
and report ONLY new or still-unaddressed issues." \
--output-format json --json-schema .claude/schemas/review-findings.json \
> findings.json
# Test generation: existing tests in context, standards from CLAUDE.md
claude -p "Generate tests for the changed files. Existing suites are in tests/. \
Do not duplicate covered scenarios. Follow the testing standards and use the \
fixtures documented in CLAUDE.md." \
--output-format json --json-schema .claude/schemas/test-plan.jsonAnti-patterns
- Invoking Claude Code in CI without
-p/--printbecause the interactive session waits for input the runner can never provide and the job hangs until it times out. - Parsing free-form prose output with regular expressions instead of using
--output-format jsonwith--json-schemabecause the shape drifts between runs and the posting step breaks silently. - Asking the same session that generated the code to review it instead of using an independent review instance because it shares the assumptions that produced the defects.
- Re-running a review on each push without supplying prior findings because every unchanged issue is reported again and the duplicate comments train the team to ignore the bot.
How it is examined
- Any stem describing a CI job that "hangs" or "waits for input" is a
-p/--printitem — pick the flag, not a timeout increase or a retry. - When the requirement is inline PR comments at specific lines, the answer combines
--output-format jsonwith--json-schema; an option offering only one of the two is usually incomplete. - Low-value or duplicate generated tests point at context fixes — CLAUDE.md standards and fixtures, plus the existing test files — not at a different model or a longer prompt.
- If the stem sets up "the session that wrote the code should also review it", that is the distractor: session context isolation says use an independent instance.
References — 4 sources
- Run Claude Code programmatically Anthropic The normative page for non-interactive runs, including the flags that cannot combine with `-p` — `--bg`, and `--cloud` with a task description — a failure a CI author hits immediately.
- CLI reference Anthropic `--output-format` is a selector with three values — `text` (the default), `json`, `stream-json` — not a switch. It also settles the argument form this page hedges on: `--json-schema` takes the schema itself as an inline JSON string, is print-mode only, and exits with an error on an invalid schema.
- Get structured output from agents Anthropic Where the validated result actually lands — the `structured_output` field on the result message — plus the JSON Schema, Zod and Pydantic paths a pipeline step parses.
- Code Review Anthropic The shipped implementation of this section: parallel specialized agents, a verification step that "checks candidates against actual code behavior to filter out false positives", dedup, and inline posting — with a `REVIEW.md` injected into every review agent as the criteria artifact.
Live product docs — where they differ from the exam guide, answer from the guide. All references
Exam guide, verbatim — what is measured
Knowledge of
- The -p (or --print) flag for running Claude Code in non-interactive mode in automated pipelines
- --output-format json and --json-schema CLI flags for enforcing structured output in CI contexts
- CLAUDE.md as the mechanism for providing project context (testing standards, fixture conventions, review criteria) to CI-invoked Claude Code
- Session context isolation: why the same Claude session that generated code is less effective at reviewing its own changes compared to an independent review instance
Skills in
- Running Claude Code in CI with the -p flag to prevent interactive input hangs
- Using --output-format json with --json-schema to produce machine-parseable structured findings for automated posting as inline PR comments
- Including prior review findings in context when re-running reviews after new commits, instructing Claude to report only new or still-unaddressed issues to avoid duplicate comments
- Providing existing test files in context so test generation avoids suggesting duplicate scenarios already covered by the test suite
- Documenting testing standards, valuable test criteria, and available fixtures in CLAUDE.md to improve test generation quality and reduce low-value test output