Files
claude-plugin/plugins/reviews/skills/audit-terraform
mroberts 37fc3fb291 Fix audit tool bootstrap and add per-run preflight
audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
  under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
  since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
  and --user-only (no system package managers, no sudo).

log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.

audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).

Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
2026-09-22 15:21:28 -05:00
..

audit-terraform

Automated, tool-driven audit of terraform / terragrunt changes.

A mechanical Python collection script builds a manifest of the change (plan output, diff-touched resources, module graph, scanner findings), slices that manifest per reviewer, and the skill fans out four parallel LLM subagents against those slices. Output is a walkthrough of what the change does plus findings grouped by severity.

Ships in the reviews plugin of the mroberts marketplace, alongside review-pr. The two are different tools: review-pr builds a guided briefing so a human can read a PR; audit-terraform runs the scanners and the review agents itself.

When it runs

Auto-activates on phrasing like "audit this terraform", "check this terragrunt change", or the explicit /audit-terraform invocation.

Modes

Invocation Mode Diff target Output
/audit-terraform local working tree vs base interactive walkthrough in chat
/audit-terraform <ref> ref <ref> vs base audit-terraform-<short>.md
/audit-terraform <pr_number> ref PR head ref vs base same as above

Local mode runs against the user's current repo. Ref/PR mode isolates the checkout first — a git worktree under ~/.claude/cache/audit-terraform/<short-ref>/ when the cwd is a git checkout of the repo, otherwise a fresh gh repo clone into the same path. The clone path is what makes the skill usable from jj workspaces, where gh/git autodetection fails.

The base ref is resolved automatically: origin/HEAD if the symbolic ref exists, else origin/main or origin/master, else the literal branch name. origin is fetched first so the comparison is against the remote tip rather than a stale local branch.

Pipeline

  1. Collection — scripts/collect-changes.py --repo --base --head --output-dir --mode. Everything below happens inside this one script.
  2. Diff scan — git diff --unified=0 <base>...<head> for changed .tf / .tf.json / .hcl files, then per-file changed line ranges are mapped through hcl_diff.py (python-hcl2) to the resource/module blocks they touch.
  3. Plan units — module_graph.py builds the module→callsite graph; resolve_plan_units.py maps each changed directory to the directories that can actually be planned (a changed shared module resolves to its callsites). Modules with no callsites are recorded as an error and reviewed diff-only.
  4. Plan execution — up to 8 plan units run concurrently. plan_runner.py picks terragrunt if the unit has a terragrunt.hcl, otherwise tofu, then runs <tool> init -input=false -no-color and <tool> plan -input=false -no-color. Full stdout lands in plans/<dir>.txt; plan_output.py parses it into resource addresses/actions and the Plan: N to add, N to change, N to destroy summary.
  5. First-pass scanners —
    • trivy config --quiet --format json <changed dirs>, normalized into trivy_findings (check id, title, severity, message, file, line range, resource type). A trivy failure is a recorded error, not a hard stop.
    • tflint --format json --chdir <dir> per changed terraform dir, normalized into tflint_findings (rule, severity, message, file, line range, doc link). Silently skipped if tflint is not on PATH.
  6. Catalog + context — catalog.py merges plan hits and diff hits into one entry per resource, tagged source: plan | diff | both. source_lookup.py attaches the block header, an evidence line, key attributes, and review context so agents rarely have to open source files.
  7. Reference sets — reference_set.py computes peer directories for each changed dir (region peers and same-component cross-env for the live/<env>/<region>/<component> layout) plus precomputed consistency norms.
  8. Manifest slicing — slicing.py writes a per-agent subset so each subagent only sees what it needs. AWS-lane slices filter the catalog to aws_* types; the tf-hygiene slice omits trivy findings; the walkthrough slice omits scanner findings entirely.
  9. Fan-out — four Task subagents dispatched in a single message, each given its manifest slice, REPO, and a findings-<agent>.json output path.
  10. Aggregation — the walkthrough payload is loaded separately; remaining findings are deduplicated on {resource, control} (first wins, loser recorded in also_flagged_by) and grouped by severity.
  11. Telemetry — one scripts/log-run.py call appends a row per agent.
  12. Output — chat message in local mode, markdown report in ref mode.

Exit behaviour

collect-changes.py returns non-zero if the git diff fails or any plan unit failed to plan. On non-zero, the skill reports manifest.json errors[] verbatim and does not run subagents. Zero changed HCL files is exit 0 with a no terraform/hcl files changed error entry.

External tools

Shelled out to by the scripts:

Tool Used by Required
git diff scanning, base-ref resolution yes
tofu init/plan for non-terragrunt units yes, if any plan unit is plain terraform
terragrunt init/plan for units with a terragrunt.hcl yes, if any plan unit is terragrunt
trivy trivy config first-pass scan yes — a failure is recorded as an error
tflint per-dir lint optional — skipped if absent

gh is used by the skill procedure (not the scripts) to resolve the repo name-with-owner, look up PR head refs, and clone in ref/PR mode.

Note the plan runner invokes tofu specifically. There is no terraform binary fallback — install OpenTofu, or symlink tofu to terraform.

Install

Homebrew covers all of them:

brew install opentofu terragrunt trivy tflint gh

Upstream instructions: https://opentofu.org/docs/intro/install/, https://terragrunt.gruntwork.io/docs/getting-started/install/, https://trivy.dev/latest/getting-started/installation/, https://github.com/terraform-linters/tflint, https://cli.github.com/.

Python dependencies live in pyproject.toml: the tools group (python-hcl2, beautifulsoup4, requests) and the dev group (pytest). scripts/install-tools.sh runs uv sync --group tools into ${SKILL_DIR}/.venv/ and checks the native tools above; --check-only does just the check.

Subagents

Prompts live in agents/. Four are the default set.

Agent Model Scope
walkthrough-reviewer Sonnet Reviewer-facing summary: 3–6 sentence overview plus a per-plan-unit summary classified substantive or trivial. Describes, does not grade. Emits no findings — its output JSON is {overview, plan_units[]}.
aws-bp-reviewer Sonnet Primary security lane. Reads trivy_findings first, triages and suppresses duplicates/low-signal hits, then adds only contextual AWS best-practice findings trivy is likely to miss (encryption tradeoffs, retention/lifecycle, backup posture, multi-AZ, IAM least-privilege nuance, logging blind spots, deletion protection).
consistency-reviewer Sonnet Repo-internal only. Compares each changed dir against its reference set and precomputed norms — missing patterns peers all use, naming/variable drift, missing standard tags, missing kms_key_arn. Evidence must cite at least two peer dirs plus the diverging file:line.
tf-hygiene-reviewer Haiku 4.5 Module hygiene and maintainability, strictly non-security. Consumes tflint_findings, then adds what tflint can't infer: variable/output type and description, sensitive flags, version pinning, lifecycle blocks, terragrunt patterns. Runs on Haiku because tflint already did the mechanical work.

Legacy reviewers

agents/fsbp-reviewer.md and agents/cis-reviewer.md are benchmark-specific prompts kept for explicit fallback or cross-check work. They are not part of the default path — aws-bp-reviewer replaced them once trivy became the first-pass detector. Use them only when asked to confirm findings against FSBP or CIS specifically.

collect-changes.py still writes manifest-fsbp.json and manifest-cis.json slices so those prompts can be run without re-collecting.

Findings contract

Every finding-emitting agent writes:

{
  "agent": "aws-bp-reviewer",
  "started_at": "...",
  "finished_at": "...",
  "skipped_resources": [{"address": "...", "reason": "not AWS"}],
  "findings": [
    {
      "resource": "aws_s3_bucket.audit_logs",
      "dir": "live/prod/us-east-1/audit",
      "control": "TRIVY AVD-AWS-0089",
      "severity": "critical | high | medium | low",
      "issue": "...",
      "evidence": "main.tf:21 — ...",
      "fix": "..."
    }
  ]
}

Controls data

data/controls/ holds JSON snapshots of AWS security control catalogues, indexed by terraform resource type:

  • fsbp.json — AWS Foundational Security Best Practices (368 controls in the committed snapshot)
  • cis.json — CIS AWS Foundations Benchmark (65 controls)
  • meta.json — source URL and fetch timestamp per file

Refresh from the AWS Security Hub docs:

uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/refresh-controls.py [--output-dir DIR]

refresh-controls.py fetches the FSBP standard index and the CIS benchmark page over HTTPS, parses them with scrape_fsbp.py / scrape_cis.py (BeautifulSoup), validates rows through controls_schema.py, and rewrites all three files. It prints the control counts it wrote. --output-dir defaults to data/controls/ inside the skill.

The cache is treated as stale past 30 days; the benchmark agents check meta.json and fall back to a live fetch if a file is missing, empty, or stale. Only the legacy fsbp-reviewer / cis-reviewer prompts consume this data — the default four-agent path does not read it, since trivy supplies first-pass benchmark detection.

Output artifacts

The output directory is <REPO>/.audit-terraform/ in ref mode and ~/.claude/cache/audit-terraform/local-<UTC-timestamp>/ in local mode.

File Written by Contents
manifest.json collection Full manifest — base/head refs, mode, default branch, changed source dirs, plan units, resource catalog, trivy + tflint findings, module graph, errors
manifest-<agent>.json collection Per-agent slice for fsbp, cis, aws-bp, consistency, tf-hygiene, walkthrough
trivy-findings.json collection Raw trivy config JSON
tflint-findings.json collection Raw tflint JSON, keyed {"by_dir": {...}}
reference_sets.json collection Peer dirs per changed dir (consistency agent input)
consistency_norms.json collection Precomputed norms derived from the reference sets
plans/<dir>.txt collection Full plan stdout per plan unit (. becomes root, / becomes _)
findings-<agent>.json subagents Findings, or the walkthrough payload for walkthrough-reviewer

Ref mode additionally writes <REPO>/audit-terraform-<short-ref>.md: summary table, walkthrough (substantive then trivial plan units), per-dir plan summaries, findings by severity → category → resource, consistency findings, and skipped resources. The worktree/clone path is left in place and reported so follow-up review can use it.

Telemetry

Append-only JSONL at ~/.claude/cache/audit-terraform/runs.jsonl. Two record kinds, both defined in scripts/telemetry.py:

  • subagent_run — run_id, repo, mode, agent, model, input_tokens, output_tokens, duration_ms, finding_count
  • verdict — run_id, agent, rule_id, file, line, verdict (kept | dismissed | false_positive), notes

Runs are logged with one call after the fan-out completes. The orchestrator supplies token/duration metadata per agent; log-run.py reads the finding count off disk from each findings-<agent>.json:

echo '{"aws-bp-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}}' \
  | uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
      --output-dir <OUTPUT> --run-id <hex8> --repo <REPO> \
      --mode <local|ref> --usage-json -

--log-path overrides the default log location.

Read the log back with:

uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/review_stats.py

It prints JSON with by_agent (tokens, duration, runs, verdict counts, precision = kept/total, tokens_per_kept), by_rule keyed <agent>/<rule_id>, and a distinct runs count. It takes no arguments and always reads the default log path.

Verdicts are how precision gets measured. Local mode prompts for them interactively and appends via telemetry.append_verdict. Ref mode does not collect verdicts, so precision figures reflect local runs alone.

Development

Tests live in tests/ — 106 of them, covering the diff and HCL parsers, plan unit resolution, plan output parsing, manifest and slicing shapes, scanner normalization, the controls scrapers and schema, telemetry, stats, and agent prompt invariants. One test in tests/test_cli.py is marked integration and needs real tofu / terragrunt.

Run them with:

cd ${SKILL_DIR} && uv run --group tools --group dev pytest

Test imports resolve through tests/conftest.py, which puts the skill root on sys.path.

Lint with ruff; collect-changes.py, refresh-controls.py and log-run.py are exempted from E402 because they mutate sys.path before importing sibling modules.

Layout

SKILL.md              procedure — source of truth for behaviour
agents/               subagent system prompts (4 default + 2 legacy)
data/controls/        FSBP/CIS snapshots + fetch metadata
scripts/              collection pipeline, scrapers, telemetry
tests/                pytest suite + HTML fixtures for the scrapers