Files
claude-plugin/plugins/reviews/skills/audit-code/SKILL.md
T
mroberts 37fc3fb291 Fix audit tool bootstrap and add per-run preflight
audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
  under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
  since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
  and --user-only (no system package managers, no sudo).

log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.

audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).

Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
2026-09-22 15:21:28 -05:00

320 lines
13 KiB
Markdown

---
name: audit-code
description: Review source-code changes (Python, JavaScript/TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions) by running OSS linters (bandit, ruff, eslint, tsc, mypy, opengrep, gitleaks, pip-audit, osv-scanner, SecurityCodeScan via dotnet, selene, luac, PSScriptAnalyzer, InjectionHunter, actionlint, zizmor), filtering findings to the diff, and dispatching eight parallel review agents for a PR walkthrough plus security triage, type safety, dependencies, consistency, secrets, maintainability, and GitHub Actions review. Two modes — interactive review of the local working tree, or written report for a git ref / GitHub PR. This is the automated tool-driven audit; for a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on phrasing like "audit this code", "run the linters on this change", or "/audit-code".
---
# audit-code
You review source-code changes by running the OSS SAST stack against the
repo, diff-filtering the output, and fanning out eight review subagents in
parallel against per-agent manifest slices. One of the eight is a
walkthrough agent that produces the reviewer-facing summary of what the PR
does.
**Always announce at start:** "Using audit-code to walk through the change
and audit for security, type safety, dependencies, consistency, secrets,
and maintainability."
## When to invoke
- User says "review the code" / "review this PR" / "review my changes" / etc.
- User runs `/audit-code` with or without an argument.
## Modes
| Invocation | Mode | Diff target | Output |
|--------------------------------|---------|----------------------------|---------------------------------|
| `/audit-code` | local | working tree vs origin | interactive walkthrough in chat |
| `/audit-code <ref>` | ref | `<ref>` vs origin | `audit-code-<short>.md` + chat |
| `/audit-code <pr_number>` | ref | PR head ref vs base | same as ref mode |
For ref/PR mode work in a fresh worktree. For local mode work in the user's
current repo.
## Procedure
### 0. Resolve `SKILL_DIR`
`${SKILL_DIR}` below means the absolute directory containing *this* SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
```
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
```
### 0.5 Preflight — every run
```
bash ${SKILL_DIR}/scripts/install-tools.sh --check-only
```
It installs nothing and returns in milliseconds when everything is present,
so run it every time. Each plugin version runs from its own directory with a
fresh, empty `.venv`, so expect it to fail on the first run after an update.
- **Exit 0:** continue.
- **Non-zero:** it lists what is missing. Run
`bash ${SKILL_DIR}/scripts/install-tools.sh --user-only` (Python tools into
`${SKILL_DIR}/.venv`, opengrep, user-scope PowerShell modules; never sudo),
then re-run `--check-only`.
- **Still missing** (`gitleaks`, `osv-scanner`, `gh` are system packages):
ask the user before running `bash ${SKILL_DIR}/scripts/install-tools.sh`
without `--user-only` — it installs through brew/paru/yay/pacman and may
call sudo. If they decline, continue: the collection script records each
missing tool in `tools_unavailable` with its install command.
### 1. Resolve mode, repo identity, and worktree
First resolve `NWO` (owner/repo) so every `gh` call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
```fish
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
```
Always pass `--repo "$NWO"` on `gh` calls — do not let `gh` autodetect from
the cwd, since that runs `git` internally and fails in jj-only workspaces
with `fatal: not a git repository`.
Then resolve mode:
- No argument: mode = `local`, `REPO` = cwd.
- Argument matches `^[0-9]+$`: GitHub PR number. Resolve via
`gh pr view <PR> --repo "$NWO" --json headRefName,baseRefName`. If `gh`
is missing, stop with: "Install `gh` and run `gh auth login`, or pass a
git ref instead." If `NWO` is empty, stop with: "Cannot determine GitHub
repo from this directory. Run from inside a checkout of the repo, or
pass an explicit ref."
- Any other string: git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being a
git repo:
- If a native worktree tool (e.g. `EnterWorktree`) is available AND the cwd
is a git checkout of `$NWO`, prefer `superpowers:using-git-worktrees` to
create `~/.claude/cache/audit-code/<short-ref>/`.
- Otherwise (jj workspaces, or invoked from outside the repo), do a fresh
clone — this never touches the surrounding workspace:
```
gh repo clone "$NWO" ~/.claude/cache/audit-code/<short-ref>/
git -C ~/.claude/cache/audit-code/<short-ref>/ checkout <ref>
```
For PR mode, `<ref>` is the head branch returned by `gh pr view` above.
> **jj users:** Local mode works in jj workspaces that have a colocated
> `.git` (the script reads the working tree but resolves the diff base via
> git). For non-colocated additional `jj workspace add` checkouts, run
> ref/PR mode instead — the fresh-clone path works regardless of cwd.
### 2. Resolve output directory
- Ref mode: `OUTPUT = <REPO>/.audit-code/`
- Local mode: `OUTPUT = ~/.claude/cache/audit-code/local-<UTC-timestamp>/`
Create the directory.
### 3. Run the collection script
```
uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/collect-findings.py \
--repo <REPO> --head <head-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
```
- If exit code != 0, read `<OUTPUT>/manifest.json` `errors[]` and report
them verbatim. Stop.
- The script handles the origin-base fetch, language detection, linter
dispatch, diff filtering, and writes manifest + per-agent slices.
### 4. Fan out eight subagents IN PARALLEL
Dispatch eight Task subagents in a single message. Each gets the system
prompt from `${SKILL_DIR}/agents/<agent>.md` and these
variables substituted:
- `MANIFEST = <OUTPUT>/manifest-<agent>.json`
- `REPO = <REPO>`
- `MODE = <local|ref>`
- `OUTPUT = <OUTPUT>/findings-<agent>.json` (walkthrough uses the same
filename but its payload is the walkthrough JSON, not findings).
Agents: `walkthrough-reviewer`, `security-triage-reviewer`,
`type-safety-reviewer`, `dependency-reviewer`, `consistency-reviewer`,
`secrets-reviewer`, `maintainability-reviewer`, `gha-reviewer`.
`walkthrough-reviewer` produces the reviewer-facing PR summary — overview
plus adaptive per-file walkthrough. It emits no findings and does no
triage.
`maintainability-reviewer` runs on Haiku 4.5 by default — its work is
deterministic-tool triage and doesn't benefit from a larger model.
Override with `--maintainability-model`.
### 5. Aggregate findings
After all eight return:
1. Load `findings-walkthrough-reviewer.json` separately as the
`walkthrough` payload (`overview` + `files[]`). Discard if malformed
and note in summary.
2. Load each `findings-<agent>.json` for the remaining seven. Discard
malformed files (note in summary).
3. Deduplicate findings sharing `{file, line, rule_id}`. First-to-finish
wins; add `also_flagged_by` with the loser's agent + rule_id.
4. Group findings by severity: critical, high, medium, low, info.
### 5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
**once**. It scans `<OUTPUT>/findings-<agent>.json` for each agent and
appends one `subagent_run` row to `~/.claude/cache/audit-code/runs.jsonl`
with the finding count read from the file, plus the token / duration
metadata you pass in.
Build a JSON blob from each Task call's `<usage>` block and pipe it in:
```fish
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"security-triage-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"type-safety-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"dependency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"secrets-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"maintainability-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N},
"gha-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
```
`RUN_ID` is a short random hex (8 chars) generated once at the start of
the run; record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not
skip it.
**Per-agent default models** (used for `AGENT_MODEL` and as defaults if no
override flag is supplied):
| Agent | Default model | Override flag |
|---|---|---|
| `security-triage-reviewer` | Sonnet | `--security-model` |
| `type-safety-reviewer` | Sonnet | `--type-safety-model` |
| `dependency-reviewer` | Sonnet | `--dependency-model` |
| `consistency-reviewer` | Sonnet | `--consistency-model` |
| `secrets-reviewer` | Sonnet | `--secrets-model` |
| `maintainability-reviewer` | Haiku 4.5 | `--maintainability-model` |
| `walkthrough-reviewer` | Sonnet | `--walkthrough-model` |
| `gha-reviewer` | Sonnet | `--gha-model` |
### 6. Output
#### Local mode (interactive)
Send a chat message structured like:
```
Reviewed N files (X py, Y ts). N findings (n critical, n high, n medium, n low).
Tools: bandit ruff eslint opengrep ... ✓ | mypy ✗ (not on PATH)
WALKTHROUGH
<overview paragraph>
Substantive changes:
- <path> — <one-to-three sentences>
- ...
Trivial:
- <path> — <one-liner>
- ...
CRITICAL (n)
- <file>:<line> (<rule_id> / <CWE>) — <issue>
fix:
<code>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
DEPENDENCY (n)
- <manifest>: <package> <from> → <to> ...
CONSISTENCY (n)
- <file>: <issue>
MAINTAINABILITY (n)
- <file>:<line> — <issue>
Full findings: <OUTPUT>/findings-*.json
Ask me to drill in, expand a section, or generate fixes as patches.
```
#### Ref mode (report)
Write `<REPO>/audit-code-<short-ref>.md` with:
1. Headline table — files reviewed, findings by severity, tools used / unavailable
2. **Walkthrough** — overview paragraph, then two subsections:
- "Substantive changes" (each substantive file as a subheading with the
summary beneath)
- "Trivial changes" (one-line bullets, collapsible if long)
3. Per-file findings grouped by severity, each with `question:` framings
4. Dependency changes with CVE deltas
5. Consistency observations (separate section)
6. Maintainability observations (dead code, complexity, duplication)
7. Skipped files (unsupported languages)
8. Linter coverage (which tools ran)
Chat gets a 5-line headline + file path.
### 7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed the findings, ask for verdicts so the skill
can measure precision over time. This applies to local mode only — prompt
interactively (skip on `--no-feedback`).
For each finding, capture `{kept | dismissed | false_positive}` plus an
optional note. Append a `verdict` record per finding to
`~/.claude/cache/audit-code/runs.jsonl` via
`scripts.telemetry.append_verdict`.
> Ref-mode reports do not collect verdicts. `review_stats.py` takes no
> arguments; it only aggregates what local mode already wrote.
### 8. Always surface the worktree path
In ref mode, end with: "Worktree left at `<REPO>` for follow-up review."
## Failure modes
- **No supported source files in diff:** manifest's `errors[]` populated,
script exited non-zero. Surface verbatim.
- **Unsupported source language(s) in diff:** the script keeps reviewing the
supported files but appends an `errors[]` entry per language, of the form
`unsupported language not reviewed: <Language> (<paths>)`. Surface these
prominently at the top of the chat summary / report so the user knows
coverage was partial. Example: `⚠️ Go file changed but not reviewed:
cmd/server.go (Go support not in this skill yet).`
- **No `gh`:** see step 1.
- **Tools missing on PATH:** non-fatal; `tools_unavailable` maps each to
its install command.
Agents acknowledge degraded coverage.
- **PowerShell files changed but no `pwsh`:** `psscriptanalyzer` and
`injectionhunter` report `not on PATH` / `<module> not installed`. Both
come from `scripts/install-tools.sh`; InjectionHunter is the injection SAST
pass, so its absence means PowerShell injection flaws went unchecked — say
so rather than reporting the change as clean.
- **One agent fails:** report the other six and note the agent that failed.