Files
claude-plugin/plugins/reviews/skills/audit-code/SKILL.md
T
mroberts 37fc3fb291 Fix audit tool bootstrap and add per-run preflight
audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
  under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
  since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
  and --user-only (no system package managers, no sudo).

log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.

audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).

Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
2026-09-22 15:21:28 -05:00

13 KiB

name, description
name description
audit-code Review source-code changes (Python, JavaScript/TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions) by running OSS linters (bandit, ruff, eslint, tsc, mypy, opengrep, gitleaks, pip-audit, osv-scanner, SecurityCodeScan via dotnet, selene, luac, PSScriptAnalyzer, InjectionHunter, actionlint, zizmor), filtering findings to the diff, and dispatching eight parallel review agents for a PR walkthrough plus security triage, type safety, dependencies, consistency, secrets, maintainability, and GitHub Actions review. Two modes — interactive review of the local working tree, or written report for a git ref / GitHub PR. This is the automated tool-driven audit; for a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on phrasing like "audit this code", "run the linters on this change", or "/audit-code".

audit-code

You review source-code changes by running the OSS SAST stack against the repo, diff-filtering the output, and fanning out eight review subagents in parallel against per-agent manifest slices. One of the eight is a walkthrough agent that produces the reviewer-facing summary of what the PR does.

Always announce at start: "Using audit-code to walk through the change and audit for security, type safety, dependencies, consistency, secrets, and maintainability."

When to invoke

  • User says "review the code" / "review this PR" / "review my changes" / etc.
  • User runs /audit-code with or without an argument.

Modes

Invocation Mode Diff target Output
/audit-code local working tree vs origin interactive walkthrough in chat
/audit-code <ref> ref <ref> vs origin audit-code-<short>.md + chat
/audit-code <pr_number> ref PR head ref vs base same as ref mode

For ref/PR mode work in a fresh worktree. For local mode work in the user's current repo.

Procedure

0. Resolve SKILL_DIR

${SKILL_DIR} below means the absolute directory containing this SKILL.md. You were given that path when this skill loaded — export it once before any other command so the bundled scripts resolve wherever the plugin is installed:

export SKILL_DIR=<absolute path to the directory holding this SKILL.md>

0.5 Preflight — every run

bash ${SKILL_DIR}/scripts/install-tools.sh --check-only

It installs nothing and returns in milliseconds when everything is present, so run it every time. Each plugin version runs from its own directory with a fresh, empty .venv, so expect it to fail on the first run after an update.

  • Exit 0: continue.
  • Non-zero: it lists what is missing. Run bash ${SKILL_DIR}/scripts/install-tools.sh --user-only (Python tools into ${SKILL_DIR}/.venv, opengrep, user-scope PowerShell modules; never sudo), then re-run --check-only.
  • Still missing (gitleaks, osv-scanner, gh are system packages): ask the user before running bash ${SKILL_DIR}/scripts/install-tools.sh without --user-only — it installs through brew/paru/yay/pacman and may call sudo. If they decline, continue: the collection script records each missing tool in tools_unavailable with its install command.

1. Resolve mode, repo identity, and worktree

First resolve NWO (owner/repo) so every gh call works regardless of cwd — git checkout, jj workspace, or outside any repo:

set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
         or jj git remote list 2>/dev/null \
              | awk '$1=="origin"{print $2}' \
              | sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
         or git remote get-url origin 2>/dev/null \
              | sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')

Always pass --repo "$NWO" on gh calls — do not let gh autodetect from the cwd, since that runs git internally and fails in jj-only workspaces with fatal: not a git repository.

Then resolve mode:

  • No argument: mode = local, REPO = cwd.
  • Argument matches ^[0-9]+$: GitHub PR number. Resolve via gh pr view <PR> --repo "$NWO" --json headRefName,baseRefName. If gh is missing, stop with: "Install gh and run gh auth login, or pass a git ref instead." If NWO is empty, stop with: "Cannot determine GitHub repo from this directory. Run from inside a checkout of the repo, or pass an explicit ref."
  • Any other string: git ref.

For ref/PR mode, isolate the checkout without depending on the cwd being a git repo:

  • If a native worktree tool (e.g. EnterWorktree) is available AND the cwd is a git checkout of $NWO, prefer superpowers:using-git-worktrees to create ~/.claude/cache/audit-code/<short-ref>/.

  • Otherwise (jj workspaces, or invoked from outside the repo), do a fresh clone — this never touches the surrounding workspace:

    gh repo clone "$NWO" ~/.claude/cache/audit-code/<short-ref>/
    git -C ~/.claude/cache/audit-code/<short-ref>/ checkout <ref>
    

    For PR mode, <ref> is the head branch returned by gh pr view above.

jj users: Local mode works in jj workspaces that have a colocated .git (the script reads the working tree but resolves the diff base via git). For non-colocated additional jj workspace add checkouts, run ref/PR mode instead — the fresh-clone path works regardless of cwd.

2. Resolve output directory

  • Ref mode: OUTPUT = <REPO>/.audit-code/
  • Local mode: OUTPUT = ~/.claude/cache/audit-code/local-<UTC-timestamp>/

Create the directory.

3. Run the collection script

uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/collect-findings.py \
  --repo <REPO> --head <head-or-HEAD> \
  --output-dir <OUTPUT> --mode <local|ref>
  • If exit code != 0, read <OUTPUT>/manifest.json errors[] and report them verbatim. Stop.
  • The script handles the origin-base fetch, language detection, linter dispatch, diff filtering, and writes manifest + per-agent slices.

4. Fan out eight subagents IN PARALLEL

Dispatch eight Task subagents in a single message. Each gets the system prompt from ${SKILL_DIR}/agents/<agent>.md and these variables substituted:

  • MANIFEST = <OUTPUT>/manifest-<agent>.json
  • REPO = <REPO>
  • MODE = <local|ref>
  • OUTPUT = <OUTPUT>/findings-<agent>.json (walkthrough uses the same filename but its payload is the walkthrough JSON, not findings).

Agents: walkthrough-reviewer, security-triage-reviewer, type-safety-reviewer, dependency-reviewer, consistency-reviewer, secrets-reviewer, maintainability-reviewer, gha-reviewer.

walkthrough-reviewer produces the reviewer-facing PR summary — overview plus adaptive per-file walkthrough. It emits no findings and does no triage.

maintainability-reviewer runs on Haiku 4.5 by default — its work is deterministic-tool triage and doesn't benefit from a larger model. Override with --maintainability-model.

5. Aggregate findings

After all eight return:

  1. Load findings-walkthrough-reviewer.json separately as the walkthrough payload (overview + files[]). Discard if malformed and note in summary.
  2. Load each findings-<agent>.json for the remaining seven. Discard malformed files (note in summary).
  3. Deduplicate findings sharing {file, line, rule_id}. First-to-finish wins; add also_flagged_by with the loser's agent + rule_id.
  4. Group findings by severity: critical, high, medium, low, info.

5.5 Record telemetry — MANDATORY, one CLI call

After all subagents return (success or failure), invoke the logger once. It scans <OUTPUT>/findings-<agent>.json for each agent and appends one subagent_run row to ~/.claude/cache/audit-code/runs.jsonl with the finding count read from the file, plus the token / duration metadata you pass in.

Build a JSON blob from each Task call's <usage> block and pipe it in:

echo '{
  "walkthrough-reviewer":     {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "security-triage-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "type-safety-reviewer":     {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "dependency-reviewer":      {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "consistency-reviewer":     {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "secrets-reviewer":         {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "maintainability-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N},
  "gha-reviewer":             {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
       --output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
       --mode <local|ref> --usage-json -

RUN_ID is a short random hex (8 chars) generated once at the start of the run; record it in the chat headline so the user can correlate later verdicts back to the run.

Skipping this step means no precision or tokens-per-kept data — do not skip it.

Per-agent default models (used for AGENT_MODEL and as defaults if no override flag is supplied):

Agent Default model Override flag
security-triage-reviewer Sonnet --security-model
type-safety-reviewer Sonnet --type-safety-model
dependency-reviewer Sonnet --dependency-model
consistency-reviewer Sonnet --consistency-model
secrets-reviewer Sonnet --secrets-model
maintainability-reviewer Haiku 4.5 --maintainability-model
walkthrough-reviewer Sonnet --walkthrough-model
gha-reviewer Sonnet --gha-model

6. Output

Local mode (interactive)

Send a chat message structured like:

Reviewed N files (X py, Y ts). N findings (n critical, n high, n medium, n low).
Tools: bandit ruff eslint opengrep ... ✓  |  mypy ✗ (not on PATH)

WALKTHROUGH
  <overview paragraph>

  Substantive changes:
    - <path> — <one-to-three sentences>
    - ...
  Trivial:
    - <path> — <one-liner>
    - ...

CRITICAL (n)
  - <file>:<line> (<rule_id> / <CWE>) — <issue>
    fix:
        <code>

HIGH (n)
  - ...

MEDIUM (n) — say "expand medium" to see
LOW (n)    — say "expand low" to see

DEPENDENCY (n)
  - <manifest>: <package> <from> → <to> ...

CONSISTENCY (n)
  - <file>: <issue>

MAINTAINABILITY (n)
  - <file>:<line> — <issue>

Full findings: <OUTPUT>/findings-*.json
Ask me to drill in, expand a section, or generate fixes as patches.

Ref mode (report)

Write <REPO>/audit-code-<short-ref>.md with:

  1. Headline table — files reviewed, findings by severity, tools used / unavailable
  2. Walkthrough — overview paragraph, then two subsections:
    • "Substantive changes" (each substantive file as a subheading with the summary beneath)
    • "Trivial changes" (one-line bullets, collapsible if long)
  3. Per-file findings grouped by severity, each with question: framings
  4. Dependency changes with CVE deltas
  5. Consistency observations (separate section)
  6. Maintainability observations (dead code, complexity, duplication)
  7. Skipped files (unsupported languages)
  8. Linter coverage (which tools ran)

Chat gets a 5-line headline + file path.

7. Collect verdicts (signal-quality feedback loop)

After the user has reviewed the findings, ask for verdicts so the skill can measure precision over time. This applies to local mode only — prompt interactively (skip on --no-feedback).

For each finding, capture {kept | dismissed | false_positive} plus an optional note. Append a verdict record per finding to ~/.claude/cache/audit-code/runs.jsonl via scripts.telemetry.append_verdict.

Ref-mode reports do not collect verdicts. review_stats.py takes no arguments; it only aggregates what local mode already wrote.

8. Always surface the worktree path

In ref mode, end with: "Worktree left at <REPO> for follow-up review."

Failure modes

  • No supported source files in diff: manifest's errors[] populated, script exited non-zero. Surface verbatim.
  • Unsupported source language(s) in diff: the script keeps reviewing the supported files but appends an errors[] entry per language, of the form unsupported language not reviewed: <Language> (<paths>). Surface these prominently at the top of the chat summary / report so the user knows coverage was partial. Example: ⚠️ Go file changed but not reviewed: cmd/server.go (Go support not in this skill yet).
  • No gh: see step 1.
  • Tools missing on PATH: non-fatal; tools_unavailable maps each to its install command. Agents acknowledge degraded coverage.
  • PowerShell files changed but no pwsh: psscriptanalyzer and injectionhunter report not on PATH / <module> not installed. Both come from scripts/install-tools.sh; InjectionHunter is the injection SAST pass, so its absence means PowerShell injection flaws went unchecked — say so rather than reporting the change as clean.
  • One agent fails: report the other six and note the agent that failed.