Move the code and terraform audits into the reviews plugin

Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
2026-07-21 11:11:05 -05:00
parent 600c1fef86
commit f5934181ec
179 changed files with 20779 additions and 3 deletions
+297
View File
@@ -0,0 +1,297 @@
---
name: audit-code
description: Review source-code changes (Python, JavaScript/TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions) by running OSS linters (bandit, ruff, eslint, tsc, mypy, opengrep, gitleaks, pip-audit, osv-scanner, SecurityCodeScan via dotnet, selene, luac, PSScriptAnalyzer, InjectionHunter, actionlint, zizmor), filtering findings to the diff, and dispatching eight parallel review agents for a PR walkthrough plus security triage, type safety, dependencies, consistency, secrets, maintainability, and GitHub Actions review. Two modes — interactive review of the local working tree, or written report for a git ref / GitHub PR. This is the automated tool-driven audit; for a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on phrasing like "audit this code", "run the linters on this change", or "/audit-code".
---
# audit-code
You review source-code changes by running the OSS SAST stack against the
repo, diff-filtering the output, and fanning out eight review subagents in
parallel against per-agent manifest slices. One of the eight is a
walkthrough agent that produces the reviewer-facing summary of what the PR
does.
**Always announce at start:** "Using audit-code to walk through the change
and audit for security, type safety, dependencies, consistency, secrets,
and maintainability."
## When to invoke
- User says "review the code" / "review this PR" / "review my changes" / etc.
- User runs `/audit-code` with or without an argument.
## Modes
| Invocation | Mode | Diff target | Output |
|--------------------------------|---------|----------------------------|---------------------------------|
| `/audit-code` | local | working tree vs origin | interactive walkthrough in chat |
| `/audit-code <ref>` | ref | `<ref>` vs origin | `audit-code-<short>.md` + chat |
| `/audit-code <pr_number>` | ref | PR head ref vs base | same as ref mode |
For ref/PR mode work in a fresh worktree. For local mode work in the user's
current repo.
## Procedure
### 0. Resolve `SKILL_DIR`
`${SKILL_DIR}` below means the absolute directory containing *this* SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
```
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
```
### 1. Resolve mode, repo identity, and worktree
First resolve `NWO` (owner/repo) so every `gh` call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
```fish
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
```
Always pass `--repo "$NWO"` on `gh` calls — do not let `gh` autodetect from
the cwd, since that runs `git` internally and fails in jj-only workspaces
with `fatal: not a git repository`.
Then resolve mode:
- No argument: mode = `local`, `REPO` = cwd.
- Argument matches `^[0-9]+$`: GitHub PR number. Resolve via
`gh pr view <PR> --repo "$NWO" --json headRefName,baseRefName`. If `gh`
is missing, stop with: "Install `gh` and run `gh auth login`, or pass a
git ref instead." If `NWO` is empty, stop with: "Cannot determine GitHub
repo from this directory. Run from inside a checkout of the repo, or
pass an explicit ref."
- Any other string: git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being a
git repo:
- If a native worktree tool (e.g. `EnterWorktree`) is available AND the cwd
is a git checkout of `$NWO`, prefer `superpowers:using-git-worktrees` to
create `~/.claude/cache/audit-code/<short-ref>/`.
- Otherwise (jj workspaces, or invoked from outside the repo), do a fresh
clone — this never touches the surrounding workspace:
```
gh repo clone "$NWO" ~/.claude/cache/audit-code/<short-ref>/
git -C ~/.claude/cache/audit-code/<short-ref>/ checkout <ref>
```
For PR mode, `<ref>` is the head branch returned by `gh pr view` above.
> **jj users:** Local mode works in jj workspaces that have a colocated
> `.git` (the script reads the working tree but resolves the diff base via
> git). For non-colocated additional `jj workspace add` checkouts, run
> ref/PR mode instead — the fresh-clone path works regardless of cwd.
### 2. Resolve output directory
- Ref mode: `OUTPUT = <REPO>/.audit-code/`
- Local mode: `OUTPUT = ~/.claude/cache/audit-code/local-<UTC-timestamp>/`
Create the directory.
### 3. Run the collection script
```
python ${SKILL_DIR}/scripts/collect-findings.py \
--repo <REPO> --head <head-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
```
- If exit code != 0, read `<OUTPUT>/manifest.json` `errors[]` and report
them verbatim. Stop.
- The script handles the origin-base fetch, language detection, linter
dispatch, diff filtering, and writes manifest + per-agent slices.
### 4. Fan out eight subagents IN PARALLEL
Dispatch eight Task subagents in a single message. Each gets the system
prompt from `${SKILL_DIR}/agents/<agent>.md` and these
variables substituted:
- `MANIFEST = <OUTPUT>/manifest-<agent>.json`
- `REPO = <REPO>`
- `MODE = <local|ref>`
- `OUTPUT = <OUTPUT>/findings-<agent>.json` (walkthrough uses the same
filename but its payload is the walkthrough JSON, not findings).
Agents: `walkthrough-reviewer`, `security-triage-reviewer`,
`type-safety-reviewer`, `dependency-reviewer`, `consistency-reviewer`,
`secrets-reviewer`, `maintainability-reviewer`, `gha-reviewer`.
`walkthrough-reviewer` produces the reviewer-facing PR summary — overview
plus adaptive per-file walkthrough. It emits no findings and does no
triage.
`maintainability-reviewer` runs on Haiku 4.5 by default — its work is
deterministic-tool triage and doesn't benefit from a larger model.
Override with `--maintainability-model`.
### 5. Aggregate findings
After all eight return:
1. Load `findings-walkthrough-reviewer.json` separately as the
`walkthrough` payload (`overview` + `files[]`). Discard if malformed
and note in summary.
2. Load each `findings-<agent>.json` for the remaining seven. Discard
malformed files (note in summary).
3. Deduplicate findings sharing `{file, line, rule_id}`. First-to-finish
wins; add `also_flagged_by` with the loser's agent + rule_id.
4. Group findings by severity: critical, high, medium, low, info.
### 5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
**once**. It scans `<OUTPUT>/findings-<agent>.json` for each agent and
appends one `subagent_run` row to `~/.claude/cache/audit-code/runs.jsonl`
with the finding count read from the file, plus the token / duration
metadata you pass in.
Build a JSON blob from each Task call's `<usage>` block and pipe it in:
```fish
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"security-triage-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"type-safety-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"dependency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"secrets-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"maintainability-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N},
"gha-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
```
`RUN_ID` is a short random hex (8 chars) generated once at the start of
the run; record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not
skip it.
**Per-agent default models** (used for `AGENT_MODEL` and as defaults if no
override flag is supplied):
| Agent | Default model | Override flag |
|---|---|---|
| `security-triage-reviewer` | Sonnet | `--security-model` |
| `type-safety-reviewer` | Sonnet | `--type-safety-model` |
| `dependency-reviewer` | Sonnet | `--dependency-model` |
| `consistency-reviewer` | Sonnet | `--consistency-model` |
| `secrets-reviewer` | Sonnet | `--secrets-model` |
| `maintainability-reviewer` | Haiku 4.5 | `--maintainability-model` |
| `walkthrough-reviewer` | Sonnet | `--walkthrough-model` |
| `gha-reviewer` | Sonnet | `--gha-model` |
### 6. Output
#### Local mode (interactive)
Send a chat message structured like:
```
Reviewed N files (X py, Y ts). N findings (n critical, n high, n medium, n low).
Tools: bandit ruff eslint opengrep ... ✓ | mypy ✗ (not on PATH)
WALKTHROUGH
<overview paragraph>
Substantive changes:
- <path> — <one-to-three sentences>
- ...
Trivial:
- <path> — <one-liner>
- ...
CRITICAL (n)
- <file>:<line> (<rule_id> / <CWE>) — <issue>
fix:
<code>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
DEPENDENCY (n)
- <manifest>: <package> <from> → <to> ...
CONSISTENCY (n)
- <file>: <issue>
MAINTAINABILITY (n)
- <file>:<line> — <issue>
Full findings: <OUTPUT>/findings-*.json
Ask me to drill in, expand a section, or generate fixes as patches.
```
#### Ref mode (report)
Write `<REPO>/audit-code-<short-ref>.md` with:
1. Headline table — files reviewed, findings by severity, tools used / unavailable
2. **Walkthrough** — overview paragraph, then two subsections:
- "Substantive changes" (each substantive file as a subheading with the
summary beneath)
- "Trivial changes" (one-line bullets, collapsible if long)
3. Per-file findings grouped by severity, each with `question:` framings
4. Dependency changes with CVE deltas
5. Consistency observations (separate section)
6. Maintainability observations (dead code, complexity, duplication)
7. Skipped files (unsupported languages)
8. Linter coverage (which tools ran)
Chat gets a 5-line headline + file path.
### 7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed the findings, ask for verdicts so the skill
can measure precision over time. This applies to local mode only — prompt
interactively (skip on `--no-feedback`).
For each finding, capture `{kept | dismissed | false_positive}` plus an
optional note. Append a `verdict` record per finding to
`~/.claude/cache/audit-code/runs.jsonl` via
`scripts.telemetry.append_verdict`.
> Ref-mode reports do not collect verdicts. `review_stats.py` takes no
> arguments; it only aggregates what local mode already wrote.
### 8. Always surface the worktree path
In ref mode, end with: "Worktree left at `<REPO>` for follow-up review."
## Failure modes
- **No supported source files in diff:** manifest's `errors[]` populated,
script exited non-zero. Surface verbatim.
- **Unsupported source language(s) in diff:** the script keeps reviewing the
supported files but appends an `errors[]` entry per language, of the form
`unsupported language not reviewed: <Language> (<paths>)`. Surface these
prominently at the top of the chat summary / report so the user knows
coverage was partial. Example: `⚠️ Go file changed but not reviewed:
cmd/server.go (Go support not in this skill yet).`
- **No `gh`:** see step 1.
- **Tools missing on PATH:** non-fatal; `tools_unavailable` lists them.
Agents acknowledge degraded coverage.
- **PowerShell files changed but no `pwsh`:** `psscriptanalyzer` and
`injectionhunter` report `not on PATH` / `<module> not installed`. Both
come from `scripts/install-tools.sh`; InjectionHunter is the injection SAST
pass, so its absence means PowerShell injection flaws went unchecked — say
so rather than reporting the change as clean.
- **One agent fails:** report the other six and note the agent that failed.