Skip to content
CCAR-FAcademy
Domain 3 · Statement 3.6 6 of 6
3.6

Integrate Claude Code into CI/CD pipelines

  • Use -p (--print) so Claude Code runs non-interactively and the CI job never hangs waiting for input.
  • --output-format json plus --json-schema produce machine-parseable findings a pipeline can post as inline PR comments.
  • CLAUDE.md is how a CI-invoked run gets testing standards, fixture conventions and review criteria.
  • An independent review instance beats the session that wrote the code, which is biased toward its own assumptions.
  • Feed prior review findings back in and ask for only new or still-unaddressed issues, so re-runs do not duplicate comments.
  • Provide the existing test files so generated tests do not duplicate scenarios already covered.

Running Claude Code in a pipeline changes three things: there is no human to answer prompts, the output must be parsed by a machine, and the session has no accumulated project knowledge unless you give it some.

Surface by surface

What the pipeline is missing The mechanism What it produces
Nobody is there to answer a prompt the -p (or --print) flag a non-interactive run that takes the prompt, produces output and exits — the flag that prevents a CI job from hanging, and the single most-tested fact in this task statement42
A machine, not a reader, consumes the findings --output-format json with --json-schema structured output under an enforced schema — file, line, severity, message — that a pipeline step loops over and posts as inline pull-request comments at the right locations, instead of one generic paragraph44
The run knows nothing about the project CLAUDE.md, read straight from the checked-out repository testing standards, fixture conventions and review criteria on every invocation
Generated tests are trivial, or re-cover ground the suite already holds the existing test files in context, plus CLAUDE.md on what makes a test valuable and which fixtures already exist tests that add coverage and reuse the team's fixtures, rather than getter tests and hand-rolled fixture objects
The reviewer helped write what it is reviewing a separate invocation, never a continuation of the generating session an independent review instance that sees only the diff and the standards
A re-run re-reports everything the prior review findings in context, with an instruction to report only new or still-unaddressed issues a PR thread that stays readable across ten pushes

What matters for the exam is that the two output flags exist and are used together. --output-format is a selector, not a switch: it takes text (the default), json, and stream-json.43 The guide tests json because a CI step needs one parseable document at the end — stream-json emits messages as they arrive, which is for a live consumer, not a posting script. If a stem offers stream-json for a "post findings as PR comments" job, it is the plausible-but-wrong option.

The way a schema is handed to --json-schema in the snippets below — a path to a schema file — is written as an illustration, not as a quoted argument specification.43

This is the CLI/CI layer. At the API layer the equivalent mechanism is tool use with a JSON input_schema (Domain 4, task statement 4.3), which is described there as the most reliable way to guarantee schema-compliant output. Same goal — a shape the caller can rely on — enforced at two different layers, so read the stem for which one it is asking about.

Session context isolation

A subtle but heavily tested point: the same session that generated code is less effective at reviewing its own changes than an independent review instance. The generating session carries the assumptions and justifications that produced the code, so it is primed to consider them correct. A fresh instance judges the diff on its merits, which is why the review step is its own job in the pipeline.45

Noise is a context problem, not a model problem

Two of those rows are the same move made twice: hand the run the context that lets it stay quiet. Undocumented fixtures produce hand-rolled objects and trivial assertions. Missing prior findings produce a thread of repeats. Either way the team learns to ignore the bot, and that is the real cost.

Claude Code in a CI review pipelineShow the full non-interactive review loop: what context goes in (CLAUDE.md, diff, existing tests, prior findings), which flags shape the output, and how structured findings become inline PR comments without duplicates.Pull request pushPull requestpushCI job startsCI job startsCLAUDE.md standards and criteriaCLAUDE.mdstandardsand criteriaDiff and existing testsDiff andexistingtestsPrior review findingsPrior reviewfindingsclaude -p (non-interactive)claude -p(non-interactive)--output-format json --json-schema--output-formatjson--json-schemaValidated findings JSONValidatedfindingsJSONPost inline PR commentsPost inlinePRcommentswebhook triggersworkflowreview criteria andfixturesavoid duplicatescenariosreport only new orunaddressedno prompt can hangthe jobindependent reviewinstancesuppress repeatsenforce outputshapefile, line, severityone comment perfindingnext push re-enterswith prior findings
Claude Code in a CI review pipeline

Show the full non-interactive review loop: what context goes in (CLAUDE.md, diff, existing tests, prior findings), which flags shape the output, and how structured findings become inline PR comments without duplicates.

one transition at a time

Click a context input to see what noise problem it prevents.

Self-review vs independent review instanceExplain why the session that generated the code is a weaker reviewer than a fresh instance that sees only the diff and the project standards.Self-reviewindependent review instanceSession A (wrote the code)Session A (wrotethe code)Session B (independent reviewer)Session B(independentreviewer)Carries its own assumptionsCarries its ownassumptionsSees only diff plus CLAUDE.mdSees only diffplus CLAUDE.mdMisses defects it justifiedMisses defects itjustifiedJudges code on its meritsJudges code onits meritsdesign rationale still incontextprimed to consider themcorrectno prior rationalestronger defect detectionextended thinking is not asubstitute for a separatereview instance
Self-review vs independent review instance

Explain why the session that generated the code is a weaker reviewer than a fresh instance that sees only the diff and the project standards.

Automated PR review step in CI

Scenario 5 · Claude Code for Continuous Integration

The review job runs Claude Code with -p so it cannot block on input, and with --output-format json --json-schema so the result is validated structured data. A following step iterates the findings array and posts each one as an inline comment on the exact file and line.

Note what is deliberately not in this job: it is a separate invocation from anything that generated code, so the reviewer has no stake in the implementation. And the review criteria are not inlined in the prompt — they live in CLAUDE.md, where the team maintains them under review.

yaml
- name: Claude review
  run: |
    git diff origin/${{ github.base_ref }}...HEAD > /tmp/diff.patch
    claude -p "Review the diff in /tmp/diff.patch against the review criteria \
      in CLAUDE.md. Report only actionable defects with file and line." \
      --output-format json \
      --json-schema .claude/schemas/review-findings.json \
      > findings.json

- name: Post inline comments
  run: node scripts/post-inline-comments.js findings.json
.github/workflows/claude-review.yml — illustrative excerpt, not a quoted CLI specification

A schema that makes findings postable

Scenario 5 · Claude Code for Continuous Integration

Without a schema, one run returns { "issues": [...] } and the next returns markdown prose, and the posting script breaks. --json-schema fixes the shape so the pipeline can rely on it, and constraining severity to an enum lets the job fail the build only on high-severity findings — a direct lever on false-positive noise.

json
{
  "type": "object",
  "required": ["findings"],
  "properties": {
    "findings": {
      "type": "array",
      "items": {
        "type": "object",
        "required": ["file", "line", "severity", "message"],
        "properties": {
          "file":     { "type": "string" },
          "line":     { "type": "integer", "minimum": 1 },
          "severity": { "enum": ["high", "medium", "low"] },
          "message":  { "type": "string" },
          "isNew":    { "type": "boolean" }
        },
        "additionalProperties": false
      }
    }
  }
}
An example findings schema — the file name and path are arbitrary

Re-running the review after new commits, and generating tests that add value

Scenario 5 · Claude Code for Continuous Integration

Two noise problems, two context fixes.

Duplicate comments. On the second push the job passes the previous findings.json back in and instructs Claude to report only issues that are new or still unaddressed. Resolved findings disappear from the thread instead of being restated.

Low-value tests. The test-generation job supplies the existing test files so Claude can see which scenarios are already covered, and relies on CLAUDE.md for what the team considers a valuable test and which fixtures exist. Without this, the job reliably produces getter tests and hand-rolled fixture objects that duplicate tests/fixtures/.

bash
# Re-review: pass prior findings so nothing is reported twice
claude -p "Prior findings are in prior-findings.json. Review the current diff \
  and report ONLY new or still-unaddressed issues." \
  --output-format json --json-schema .claude/schemas/review-findings.json \
  > findings.json

# Test generation: existing tests in context, standards from CLAUDE.md
claude -p "Generate tests for the changed files. Existing suites are in tests/. \
  Do not duplicate covered scenarios. Follow the testing standards and use the \
  fixtures documented in CLAUDE.md." \
  --output-format json --json-schema .claude/schemas/test-plan.json
Both jobs are non-interactive and context-fed — illustrative invocations
  • Invoking Claude Code in CI without -p/--print because the interactive session waits for input the runner can never provide and the job hangs until it times out.
  • Parsing free-form prose output with regular expressions instead of using --output-format json with --json-schema because the shape drifts between runs and the posting step breaks silently.
  • Asking the same session that generated the code to review it instead of using an independent review instance because it shares the assumptions that produced the defects.
  • Re-running a review on each push without supplying prior findings because every unchanged issue is reported again and the duplicate comments train the team to ignore the bot.
  • Any stem describing a CI job that "hangs" or "waits for input" is a -p / --print item — pick the flag, not a timeout increase or a retry.
  • When the requirement is inline PR comments at specific lines, the answer combines --output-format json with --json-schema; an option offering only one of the two is usually incomplete.
  • Low-value or duplicate generated tests point at context fixes — CLAUDE.md standards and fixtures, plus the existing test files — not at a different model or a longer prompt.
  • If the stem sets up "the session that wrote the code should also review it", that is the distractor: session context isolation says use an independent instance.
References — 4 sources
  1. Run Claude Code programmatically Anthropic The normative page for non-interactive runs, including the flags that cannot combine with `-p` — `--bg`, and `--cloud` with a task description — a failure a CI author hits immediately.
  2. CLI reference Anthropic `--output-format` is a selector with three values — `text` (the default), `json`, `stream-json` — not a switch. It also settles the argument form this page hedges on: `--json-schema` takes the schema itself as an inline JSON string, is print-mode only, and exits with an error on an invalid schema.
  3. Get structured output from agents Anthropic Where the validated result actually lands — the `structured_output` field on the result message — plus the JSON Schema, Zod and Pydantic paths a pipeline step parses.
  4. Code Review Anthropic The shipped implementation of this section: parallel specialized agents, a verification step that "checks candidates against actual code behavior to filter out false positives", dedup, and inline posting — with a `REVIEW.md` injected into every review agent as the criteria artifact.
All sources verified ·

Live product docs — where they differ from the exam guide, answer from the guide. All references

Exam guide, verbatim — what is measured

Knowledge of

  • The -p (or --print) flag for running Claude Code in non-interactive mode in automated pipelines
  • --output-format json and --json-schema CLI flags for enforcing structured output in CI contexts
  • CLAUDE.md as the mechanism for providing project context (testing standards, fixture conventions, review criteria) to CI-invoked Claude Code
  • Session context isolation: why the same Claude session that generated code is less effective at reviewing its own changes compared to an independent review instance

Skills in

  • Running Claude Code in CI with the -p flag to prevent interactive input hangs
  • Using --output-format json with --json-schema to produce machine-parseable structured findings for automated posting as inline PR comments
  • Including prior review findings in context when re-running reviews after new commits, instructing Claude to report only new or still-unaddressed issues to avoid duplicate comments
  • Providing existing test files in context so test generation avoids suggesting duplicate scenarios already covered by the test suite
  • Documenting testing standards, valuable test criteria, and available fixtures in CLAUDE.md to improve test generation quality and reduce low-value test output
Back to top