Move the code and terraform audits into the reviews plugin

Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
2026-07-21 11:11:05 -05:00
parent 600c1fef86
commit f5934181ec
179 changed files with 20779 additions and 3 deletions
@@ -0,0 +1,292 @@
---
name: audit-terraform
description: Automated tool-driven audit of terraform or terragrunt changes — runs trivy/tflint/terraform validate and dispatches parallel agents for plan validity, AWS security posture, and repo consistency, in the local working tree or a git ref / PR. For a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on "audit this terraform", "check this terragrunt change", or "/audit-terraform".
---
# audit-terraform
You review terraform/terragrunt changes by running a mechanical collection
script, collecting first-pass Trivy config findings, and fanning out a
targeted LLM review pass against its manifest. One of the parallel agents
is a walkthrough agent that produces the reviewer-facing summary of what
the change does, plan-unit by plan-unit.
**Always announce at start:** "Using audit-terraform to walk through the
change and audit against plan validity, Trivy findings, AWS best practices,
and consistency."
## When to invoke
- User says "review terraform" / "review the terragrunt change" / etc.
- User runs `/audit-terraform` with or without an argument.
## Modes
| Invocation | Mode | Diff target | Output |
|-------------------------------------|---------|----------------------------|---------------------------------|
| `/audit-terraform` | local | local working tree vs base | interactive walkthrough in chat |
| `/audit-terraform <ref>` | ref | `<ref>` vs base | `audit-terraform-<short>.md` |
| `/audit-terraform <pr_number>` | ref | PR head ref vs base | same as above |
For ref/PR mode you must work in a fresh worktree. For local mode you work
in the user's current repo.
## Procedure
### 0. Resolve `SKILL_DIR`
`${SKILL_DIR}` below means the absolute directory containing *this* SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
```
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
```
### 1. Resolve mode, repo identity, and worktree
First resolve `NWO` (owner/repo) so every `gh` call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
```fish
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
```
Always pass `--repo "$NWO"` on `gh` calls — do not let `gh` autodetect from
the cwd, since that runs `git` internally and fails in jj-only workspaces
with `fatal: not a git repository`.
Then resolve mode:
- No argument: mode = `local`. `REPO = <cwd>`.
- Argument matches `^[0-9]+$`: GitHub PR number. Resolve the head ref:
```
gh pr view <PR> --repo "$NWO" --json headRefName,headRepository \
-q '.headRefName + "@" + .headRepository.url'
```
If `gh` is missing or not authenticated, stop with: "Install `gh` and run
`gh auth login`, or pass a git ref instead of a PR number." If `NWO` is
empty, stop with: "Cannot determine GitHub repo from this directory.
Run from inside a checkout of the repo, or pass an explicit ref."
- Any other string: treat as a git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being
a git repo:
- If a native worktree tool (e.g. `EnterWorktree`) is available AND the cwd
is a git checkout of `$NWO`, prefer `superpowers:using-git-worktrees` to
create `~/.claude/cache/audit-terraform/<short-ref>/`.
- Otherwise (jj workspaces, or invoked from outside the repo), do a fresh
clone — this never touches the surrounding workspace:
```
gh repo clone "$NWO" ~/.claude/cache/audit-terraform/<short-ref>/
git -C ~/.claude/cache/audit-terraform/<short-ref>/ checkout <ref>
```
For PR mode, `<ref>` is the head branch returned by `gh pr view` above.
`REPO = ~/.claude/cache/audit-terraform/<short-ref>/`.
> **jj users:** Local mode works in jj workspaces that have a colocated
> `.git` (the script reads the working tree but resolves the diff base via
> git). For non-colocated additional `jj workspace add` checkouts, run
> ref/PR mode instead — the fresh-clone path works regardless of cwd.
### 2. Resolve output directory
- Ref mode: `OUTPUT = <REPO>/.audit-terraform/`
- Local mode: `OUTPUT = ~/.claude/cache/audit-terraform/local-<UTC-timestamp>/`
Create the directory.
### 3. Run the collection script
```
python ${SKILL_DIR}/scripts/collect-changes.py \
--repo <REPO> --base <base-or-detect> --head <head-ref-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
```
- If exit code != 0, read `<OUTPUT>/manifest.json` `errors[]` and report them
to the user verbatim. Do not run subagents. Stop.
- If the manifest has zero `plan_units` AND zero `catalog` entries, tell the
user "no terraform changes detected" and stop.
- The collection step also writes raw Trivy config results to
`<OUTPUT>/trivy-findings.json` and normalizes compact `trivy_findings` into
the manifest for downstream security review.
### 4. Fan out focused subagents IN PARALLEL
Dispatch four Task subagents in a single message (parallel execution). Each
gets the system prompt from `${SKILL_DIR}/agents/<agent>.md`
and these variables substituted into its task:
- `MANIFEST = <OUTPUT>/manifest-<agent>.json` (per-agent slice; full manifest stays at `<OUTPUT>/manifest.json` for debugging)
- `REFERENCE_SETS = <OUTPUT>/reference_sets.json` (consistency only)
- `CONSISTENCY_NORMS = <OUTPUT>/consistency_norms.json` (consistency only)
- `REPO = <REPO>`
- `OUTPUT = <OUTPUT>/findings-<agent>.json` (walkthrough uses the same
filename but its payload is the walkthrough JSON, not findings).
Agents: `walkthrough-reviewer`, `aws-bp-reviewer`, `consistency-reviewer`,
`tf-hygiene-reviewer`.
Default path:
- `walkthrough-reviewer` produces the reviewer-facing PR summary —
overview plus adaptive per-plan-unit walkthrough. Emits no findings.
- `aws-bp-reviewer` consumes normalized `trivy_findings` first, suppresses
overlap/noise, and adds only contextual AWS best-practice findings Trivy is
likely to miss.
- `consistency-reviewer` stays repo-internal only.
- `tf-hygiene-reviewer` consumes `tflint_findings`, suppresses noise, and
adds module-hygiene findings tflint can't infer (var/output descriptions,
version pinning, lifecycle, terragrunt patterns). Strictly non-security.
`fsbp-reviewer` and `cis-reviewer` are legacy benchmark-specific prompts kept
for explicit fallback or cross-check work, not the default review path.
### 5. Aggregate findings
After the active agents return:
1. Load `findings-walkthrough-reviewer.json` separately as the
`walkthrough` payload (`overview` + `plan_units[]`). Discard if
malformed and note in summary.
2. Load each remaining `findings-<agent>.json`. Discard any agent file
that's malformed (note in the summary).
3. Deduplicate findings sharing `{resource, control}`. First-to-finish
wins; add a `also_flagged_by` array on the survivor with the loser's
`agent` and `control`.
4. Group findings by severity: critical, high, medium, low.
### 5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
**once**. It scans `<OUTPUT>/findings-<agent>.json` for each agent and
appends one `subagent_run` row to
`~/.claude/cache/audit-terraform/runs.jsonl` with the finding count read
from the file, plus the token / duration metadata you pass in.
Build a JSON blob from each Task call's `<usage>` block and pipe it in:
```fish
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"aws-bp-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"tf-hygiene-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
```
`RUN_ID` is a short random hex (8 chars) generated once at the start of
the run. Record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not
skip it.
**Per-agent default models:**
| Agent | Default model |
|---|---|
| `walkthrough-reviewer`| Sonnet |
| `aws-bp-reviewer` | Sonnet |
| `consistency-reviewer`| Sonnet |
| `tf-hygiene-reviewer` | Haiku 4.5 |
`tf-hygiene-reviewer` runs on Haiku because tflint already did the heavy
mechanical work — the agent's job is triage + a small number of
contextual additions.
### 6. Output
#### Local mode (interactive)
Send a chat message structured like:
```
Reviewed N plan units, M resources (X plan+diff, Y diff-only, Z plan-only).
WALKTHROUGH
<overview paragraph>
Substantive plan units:
- <plan_dir> — <2-4 sentences: what + why + destroy/replace notes>
- ...
Trivial:
- <plan_dir> — <one-liner>
- ...
CRITICAL (n)
- <resource> (<dir>): <control> — <issue>
fix: <fix>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
CONSISTENCY (n)
- <dir>: <issue>
Full findings: <OUTPUT>/findings-*.json
```
Offer follow-ups: "Ask me to drill into anything, expand a section, or
generate fixes."
#### Ref mode (report)
Write `<REPO>/audit-terraform-<short-ref>.md` structured as:
1. Summary table (resources by detection source, findings by severity).
2. **Walkthrough** — overview paragraph, then two subsections:
- "Substantive plan units" (each substantive plan_dir as a subheading
with the summary beneath, destroy/replace call-outs bolded)
- "Trivial plan units" (one-line bullets)
3. Plan summary per changed dir (add/change/destroy counts).
4. Findings grouped by severity → category → resource, each with control
reference, evidence (file:line), suggested fix.
5. Consistency findings (separate section).
6. Skipped resources (non-AWS, plan-failed-but-still-reviewed, etc.).
Send a 5-line headline to chat plus the file path.
### 7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed findings, ask for verdicts so the skill can
measure precision over time. This applies to local mode only — prompt
interactively (skip on `--no-feedback`).
For each finding, capture `{kept | dismissed | false_positive}` plus an
optional note. Append a `verdict` record per finding to
`~/.claude/cache/audit-terraform/runs.jsonl` via
`scripts.telemetry.append_verdict`.
> Ref-mode reports do not collect verdicts. `review_stats.py` takes no
> arguments; it only aggregates what local mode already wrote.
### 8. Always surface the worktree path
In ref mode, end with: "Worktree left at `<REPO>` for follow-up review."
## Failure modes
- **Plan failed in some dir:** Manifest's `errors[]` populated, script exited
non-zero. Show the errors. Do NOT run agents. Tell user to fix and re-run.
- **No `gh`:** see step 1.
- **Module with no callsites:** Manifest has a warning in `errors[]`; report
it but still run the subagents (they handle diff-only entries fine).
- **One agent fails:** Report what the surviving agents found and note the
agent that failed.