Files
claude-plugin/plugins/reviews/skills/audit-code
mroberts f5934181ec Move the code and terraform audits into the reviews plugin
Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
2026-07-21 11:11:05 -05:00
..

audit-code

Automated, linter-driven audit of source-code changes. Runs an OSS SAST and quality stack against a repo, filters every finding down to lines the change actually touched, slices the result per reviewer concern, and fans out eight subagents in parallel to triage.

Part of the reviews plugin (marketplace mroberts). The sibling skill review-pr does the opposite job — a guided walkthrough for a human reading a PR. audit-code is the tool-driven one; it produces findings, not a reading plan.

Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions workflows.

When it runs

Auto-activates on phrasing like "audit this code", "run the linters on this change", or an explicit /audit-code.

Two modes:

Invocation Mode Diff target Output
/audit-code local working tree vs origin interactive walkthrough in chat
/audit-code <ref> ref <ref> vs origin audit-code-<short>.md + chat
/audit-code <pr_number> ref PR head ref vs base same as ref mode

Local mode runs in the user's current checkout. Ref/PR mode isolates the checkout first — a git worktree when the cwd is a git checkout of the target repo, otherwise a fresh gh repo clone into ~/.claude/cache/audit-code/<short-ref>/. The clone path is what makes the skill work from jj workspaces and from outside the repo entirely.

PR-number resolution requires gh on PATH and authenticated. The repo is resolved to owner/repo up front and passed as --repo on every gh call, so gh never autodetects from cwd.

Pipeline

SKILL.md (orchestrator)
  │
  ├─► scripts/collect-findings.py --repo --head --output-dir --mode
  │      ├── resolve default branch, fetch origin, compute base
  │      ├── git diff --unified=0 base...head  → changed files + added-line ranges
  │      ├── classify each path: supported source / GHA / dep manifest / skipped
  │      ├── dispatch linters for the languages actually present
  │      ├── diff-filter: drop findings not overlapping an added-line range
  │      ├── write manifest.json
  │      └── write 8 per-agent manifest slices
  │
  ├─► 8 Task subagents IN PARALLEL, one per slice
  │      each writes findings-<agent>.json
  │
  ├─► aggregate: dedupe on {file, line, rule_id}, group by severity
  ├─► scripts/log-run.py  → append telemetry rows
  └─► chat summary (local) or markdown report (ref)

collect-findings.py exits non-zero when the diff cannot be computed, or when the change set contains no supported source files and no dependency manifests. On failure it still writes manifest.json; the errors[] array holds the reason.

Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30 others — see scripts/language_detect.py) do not abort the run. Supported files are still reviewed, and one errors[] entry per language records the gap so partial coverage is visible rather than silent.

Diff filtering

scripts/diff_filter.py keeps a finding only if its [line, end_line] span overlaps one of the added-line ranges recorded for that file. Paths are normalized (absolute → repo-relative, ./ stripped) before comparison, since adapters emit paths in whichever form their tool produced. Findings in untouched code are discarded, so the audit is scoped to the change.

Tool resolution

scripts/runner.py resolves each binary against .venv/bin/ inside the skill directory first, then falls back to PATH. A tool that resolves nowhere is recorded as ran: false with reason not on PATH and the run continues — missing tools degrade coverage, they do not fail the audit. Every tool invocation has a 180-second timeout.

External tools

All of these are shelled out to. Nothing is vendored.

Tool Fires when Purpose
bandit Python in diff Python security
ruff Python in diff Python lint (incl. S-rules)
ruff (idiom pass) Python in diff second ruff run, --isolated --select SIM,PERF,UP,RET,PLR,C90,B; reported as tool ruff-idiom
mypy Python in diff Python types
eslint JS/TS in diff JS/TS lint + security plugins
tsc JS/TS in diff TypeScript types (--noEmit)
dotnet C# in diff dotnet build --no-incremental; surfaces Roslyn + SecurityCodeScan (SCS*)
selene Lua in diff Lua lint
luac Lua in diff luac -p per file; syntax/parse errors
pwsh + PSScriptAnalyzer PowerShell in diff PowerShell lint; security rules split out
pwsh + InjectionHunter PowerShell in diff PowerShell injection SAST
actionlint .github/workflows/ or .github/actions/ YAML in diff workflow correctness
zizmor same as above workflow security (SARIF output)
opengrep always polyglot SAST, --config=auto plus bundled scripts/rules/lua-security.yaml
gitleaks always secret detection
pip-audit always Python dependency CVEs
osv-scanner always multi-ecosystem dependency CVEs
lizard always multi-language cyclomatic complexity
vulture Python in diff dead code
radon Python in diff complexity
interrogate Python in diff docstring coverage
knip (via npx) JS/TS in diff unused exports/files
jscpd JS/TS in diff copy-paste duplication

Dependency manifests (requirements*.txt, package.json, lockfiles, *.csproj, *.sln, pyproject.toml, …) are detected separately from source files. requirements*.txt and package.json additionally get a structured before/after package diff (added / removed / upgraded) computed by scripts/package_diff.py.

PowerShell is worth calling out: if pwsh is missing, both psscriptanalyzer and injectionhunter report unavailable. InjectionHunter is the only injection SAST pass for PowerShell, so its absence means PowerShell injection flaws went unchecked — that is a coverage gap, not a clean result.

Installing them

scripts/install-tools.sh

What it does:

  1. uv sync --group tools into ${SKILL_DIR}/.venv/ — installs bandit, ruff, mypy, pip-audit, vulture, radon, interrogate, lizard. Falls back to pip install bandit ruff mypy pip-audit if uv is absent.
  2. Downloads a platform-matched opengrep release binary into ${SKILL_DIR}/.venv/bin/ (Linux x86_64/aarch64, macOS x86_64/arm64).
  3. Installs the PSScriptAnalyzer and InjectionHunter PowerShell modules for the current user, if pwsh is present.
  4. Installs gitleaks, osv-scanner, and gh via Homebrew, paru/yay/pacman, or prints apt-era install instructions.
  5. Prints a resolution table for each tool, then lists the per-project tools it deliberately does not install.

Not installed by the script — these must exist in the target repo or on PATH yourself: eslint (+ eslint-plugin-security), typescript/tsc, dotnet with SecurityCodeScan, knip, jscpd, selene, luac, actionlint, zizmor.

The eight subagents

Each subagent gets the prompt from agents/<name>.md plus four substituted variables: MANIFEST, REPO, MODE, OUTPUT. Each writes findings-<agent>.json into the output directory.

Agent Manifest slice Covers
walkthrough-reviewer manifest-walkthrough.json Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings.
security-triage-reviewer manifest-security-triage.json bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, SCS* from dotnet, and the PSScriptAnalyzer credential/crypto rules.
type-safety-reviewer manifest-type-safety.json mypy, tsc, non-security eslint, non-SCS dotnet. Also hunts leaked Any, unexplained type: ignore/@ts-ignore, missing hints on new public functions.
dependency-reviewer manifest-dependency.json pip-audit and osv-scanner findings, plus the package_diffs added/removed/upgraded entries with advisory lookups.
consistency-reviewer manifest-consistency.json Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the ruff-idiom findings.
secrets-reviewer manifest-secrets.json gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names.
maintainability-reviewer manifest-maintainability.json vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the C901/PLR0915 complexity rules from the ruff idiom pass.
gha-reviewer manifest-gha-reviewer.json actionlint and zizmor findings against changed workflow files.

Slicing logic lives in scripts/slicing.py, which is the authority on which tool's output reaches which agent.

Default models: Sonnet for seven agents, Haiku 4.5 for maintainability-reviewer (its work is deterministic-tool triage). Each has an override flag — --security-model, --type-safety-model, --dependency-model, --consistency-model, --secrets-model, --maintainability-model, --walkthrough-model, --gha-model.

Aggregation dedupes findings sharing {file, line, rule_id} — first to finish wins, and the loser is recorded under also_flagged_by. Severities are critical, high, medium, low, info.

Output artifacts

Output directory:

  • Ref mode: <REPO>/.audit-code/
  • Local mode: ~/.claude/cache/audit-code/local-<UTC-timestamp>/

Contents:

File Written by Contents
manifest.json collect-findings.py full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors
manifest-<slice>.json ×8 collect-findings.py per-agent subsets
findings-<agent>.json ×8 subagents triaged findings; the walkthrough agent's file holds overview + files[] instead

Ref mode additionally writes <REPO>/audit-code-<short-ref>.md — headline table, walkthrough, per-file findings by severity, dependency deltas, consistency observations, maintainability observations, skipped files, and linter coverage. Chat gets a short headline plus the path.

Per-tool stats in the manifest record pre_filter and post_filter counts, so the noise removed by diff filtering is visible per tool.

Telemetry

Append-only JSONL at ~/.claude/cache/audit-code/runs.jsonl. Two record kinds, both written by scripts/telemetry.py:

subagent_run — one row per agent per run: run_id, repo, mode, agent, model, input_tokens, output_tokens, duration_ms, finding_count.

verdict — one row per finding the user adjudicated: run_id, agent, rule_id, file, line, verdict (one of kept, dismissed, false_positive), and an optional note.

Runs are logged by a single call after the fan-out completes:

echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
  | python ${SKILL_DIR}/scripts/log-run.py \
      --output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
      --mode <local|ref> --usage-json -

log-run.py reads token/duration metadata from the piped JSON and reads finding_count itself by counting entries in each findings-<agent>.json. --log-path overrides the default log location.

Reading it back:

uv run scripts/review_stats.py

Prints a JSON blob with runs (distinct run count), by_agent, and by_rule. Per agent: cumulative tokens, duration, run count, verdict tallies, precision (kept / total adjudicated), and tokens_per_kept. Per <agent>/<rule_id>: verdict tallies and precision. The script takes no arguments and always reads the default log path.

Verdict capture is interactive and local-mode only. Ref mode does not collect verdicts, so precision figures reflect local runs alone.

Development

uv run --group dev pytest

197 tests as of writing, covering every adapter against a recorded fixture in tests/fixtures/, plus diff filtering, git diff parsing, language detection, manifest serialization, slicing, package diffing, telemetry, log-run.py, review_stats.py, and the collect-findings.py CLI.

Tests marked integration require the real linters on PATH; the marker is declared in pyproject.toml.

Layout:

SKILL.md                    orchestration procedure (source of truth)
agents/*.md                 8 subagent prompts
scripts/collect-findings.py CLI entry point
scripts/runner.py           subprocess wrapper, tool resolution, timeouts
scripts/git_diff.py         base resolution, changed files + added-line ranges
scripts/diff_filter.py      findings → changed lines only
scripts/language_detect.py  extension → language, dep manifests, GHA paths
scripts/manifest.py         manifest dataclasses / JSON schema
scripts/slicing.py          per-agent manifest subsets
scripts/package_diff.py     requirements.txt and package.json deltas
scripts/adapters/*.py       one parser per linter → Finding[]
scripts/rules/              bundled opengrep rules
scripts/telemetry.py        JSONL append helpers
scripts/log-run.py          post-fan-out telemetry CLI
scripts/review_stats.py     telemetry aggregation CLI
scripts/install-tools.sh    tool installer
tests/                      pytest suite + recorded tool-output fixtures

Adding a linter means: write scripts/adapters/<tool>.py returning Finding[], record a fixture of its real output, add a test, wire it into the right _run_*_tools function in collect-findings.py, and add its name to the relevant tool set in scripts/slicing.py so a subagent actually receives it.