Move the code and terraform audits into the reviews plugin

Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
2026-07-21 11:11:05 -05:00
parent 600c1fef86
commit f5934181ec
179 changed files with 20779 additions and 3 deletions
@@ -0,0 +1,3 @@
__pycache__/
*.pyc
.pytest_cache/
@@ -0,0 +1,292 @@
# audit-terraform
Automated, tool-driven audit of terraform / terragrunt changes.
A mechanical Python collection script builds a manifest of the change (plan
output, diff-touched resources, module graph, scanner findings), slices that
manifest per reviewer, and the skill fans out four parallel LLM subagents
against those slices. Output is a walkthrough of what the change does plus
findings grouped by severity.
Ships in the `reviews` plugin of the `mroberts` marketplace, alongside
`review-pr`. The two are different tools: `review-pr` builds a guided
briefing so a *human* can read a PR; `audit-terraform` runs the scanners and
the review agents itself.
## When it runs
Auto-activates on phrasing like "audit this terraform", "check this
terragrunt change", or the explicit `/audit-terraform` invocation.
### Modes
| Invocation | Mode | Diff target | Output |
|--------------------------------|-------|----------------------------|---------------------------------|
| `/audit-terraform` | local | working tree vs base | interactive walkthrough in chat |
| `/audit-terraform <ref>` | ref | `<ref>` vs base | `audit-terraform-<short>.md` |
| `/audit-terraform <pr_number>` | ref | PR head ref vs base | same as above |
Local mode runs against the user's current repo. Ref/PR mode isolates the
checkout first — a git worktree under `~/.claude/cache/audit-terraform/<short-ref>/`
when the cwd is a git checkout of the repo, otherwise a fresh
`gh repo clone` into the same path. The clone path is what makes the skill
usable from jj workspaces, where `gh`/`git` autodetection fails.
The base ref is resolved automatically: `origin/HEAD` if the symbolic ref
exists, else `origin/main` or `origin/master`, else the literal branch name.
`origin` is fetched first so the comparison is against the remote tip rather
than a stale local branch.
## Pipeline
1. **Collection** — `scripts/collect-changes.py --repo --base --head
--output-dir --mode`. Everything below happens inside this one script.
2. **Diff scan** — `git diff --unified=0 <base>...<head>` for changed
`.tf` / `.tf.json` / `.hcl` files, then per-file changed line ranges are
mapped through `hcl_diff.py` (python-hcl2) to the resource/module blocks
they touch.
3. **Plan units** — `module_graph.py` builds the module→callsite graph;
`resolve_plan_units.py` maps each changed directory to the directories
that can actually be planned (a changed shared module resolves to its
callsites). Modules with no callsites are recorded as an error and
reviewed diff-only.
4. **Plan execution** — up to 8 plan units run concurrently. `plan_runner.py`
picks `terragrunt` if the unit has a `terragrunt.hcl`, otherwise `tofu`,
then runs `<tool> init -input=false -no-color` and `<tool> plan
-input=false -no-color`. Full stdout lands in `plans/<dir>.txt`;
`plan_output.py` parses it into resource addresses/actions and the
`Plan: N to add, N to change, N to destroy` summary.
5. **First-pass scanners** —
- `trivy config --quiet --format json <changed dirs>`, normalized into
`trivy_findings` (check id, title, severity, message, file, line range,
resource type). A trivy failure is a recorded error, not a hard stop.
- `tflint --format json --chdir <dir>` per changed terraform dir,
normalized into `tflint_findings` (rule, severity, message, file, line
range, doc link). Silently skipped if `tflint` is not on PATH.
6. **Catalog + context** — `catalog.py` merges plan hits and diff hits into
one entry per resource, tagged `source: plan | diff | both`.
`source_lookup.py` attaches the block header, an evidence line, key
attributes, and review context so agents rarely have to open source files.
7. **Reference sets** — `reference_set.py` computes peer directories for each
changed dir (region peers and same-component cross-env for the
`live/<env>/<region>/<component>` layout) plus precomputed consistency
norms.
8. **Manifest slicing** — `slicing.py` writes a per-agent subset so each
subagent only sees what it needs. AWS-lane slices filter the catalog to
`aws_*` types; the tf-hygiene slice omits trivy findings; the walkthrough
slice omits scanner findings entirely.
9. **Fan-out** — four Task subagents dispatched in a single message, each
given its manifest slice, `REPO`, and a `findings-<agent>.json` output
path.
10. **Aggregation** — the walkthrough payload is loaded separately; remaining
findings are deduplicated on `{resource, control}` (first wins, loser
recorded in `also_flagged_by`) and grouped by severity.
11. **Telemetry** — one `scripts/log-run.py` call appends a row per agent.
12. **Output** — chat message in local mode, markdown report in ref mode.
### Exit behaviour
`collect-changes.py` returns non-zero if the git diff fails or any plan unit
failed to plan. On non-zero, the skill reports `manifest.json` `errors[]`
verbatim and does **not** run subagents. Zero changed HCL files is exit 0
with a `no terraform/hcl files changed` error entry.
## External tools
Shelled out to by the scripts:
| Tool | Used by | Required |
|---|---|---|
| `git` | diff scanning, base-ref resolution | yes |
| `tofu` | init/plan for non-terragrunt units | yes, if any plan unit is plain terraform |
| `terragrunt` | init/plan for units with a `terragrunt.hcl` | yes, if any plan unit is terragrunt |
| `trivy` | `trivy config` first-pass scan | yes — a failure is recorded as an error |
| `tflint` | per-dir lint | optional — skipped if absent |
`gh` is used by the skill procedure (not the scripts) to resolve the repo
name-with-owner, look up PR head refs, and clone in ref/PR mode.
Note the plan runner invokes `tofu` specifically. There is no `terraform`
binary fallback — install OpenTofu, or symlink `tofu` to `terraform`.
### Install
Homebrew covers all of them:
```
brew install opentofu terragrunt trivy tflint gh
```
Upstream instructions: <https://opentofu.org/docs/intro/install/>,
<https://terragrunt.gruntwork.io/docs/getting-started/install/>,
<https://trivy.dev/latest/getting-started/installation/>,
<https://github.com/terraform-linters/tflint>, <https://cli.github.com/>.
Python dependencies are in `requirements.txt`: `python-hcl2`,
`beautifulsoup4`, `requests`, `pytest`.
## Subagents
Prompts live in `agents/`. Four are the default set.
| Agent | Model | Scope |
|---|---|---|
| `walkthrough-reviewer` | Sonnet | Reviewer-facing summary: 3–6 sentence overview plus a per-plan-unit summary classified `substantive` or `trivial`. Describes, does not grade. Emits **no** findings — its output JSON is `{overview, plan_units[]}`. |
| `aws-bp-reviewer` | Sonnet | Primary security lane. Reads `trivy_findings` first, triages and suppresses duplicates/low-signal hits, then adds only contextual AWS best-practice findings trivy is likely to miss (encryption tradeoffs, retention/lifecycle, backup posture, multi-AZ, IAM least-privilege nuance, logging blind spots, deletion protection). |
| `consistency-reviewer` | Sonnet | Repo-internal only. Compares each changed dir against its reference set and precomputed norms — missing patterns peers all use, naming/variable drift, missing standard tags, missing `kms_key_arn`. Evidence must cite at least two peer dirs plus the diverging file:line. |
| `tf-hygiene-reviewer` | Haiku 4.5 | Module hygiene and maintainability, strictly non-security. Consumes `tflint_findings`, then adds what tflint can't infer: variable/output `type` and `description`, `sensitive` flags, version pinning, lifecycle blocks, terragrunt patterns. Runs on Haiku because tflint already did the mechanical work. |
### Legacy reviewers
`agents/fsbp-reviewer.md` and `agents/cis-reviewer.md` are benchmark-specific
prompts kept for explicit fallback or cross-check work. They are **not** part
of the default path — `aws-bp-reviewer` replaced them once trivy became the
first-pass detector. Use them only when asked to confirm findings against
FSBP or CIS specifically.
`collect-changes.py` still writes `manifest-fsbp.json` and
`manifest-cis.json` slices so those prompts can be run without re-collecting.
### Findings contract
Every finding-emitting agent writes:
```json
{
"agent": "aws-bp-reviewer",
"started_at": "...",
"finished_at": "...",
"skipped_resources": [{"address": "...", "reason": "not AWS"}],
"findings": [
{
"resource": "aws_s3_bucket.audit_logs",
"dir": "live/prod/us-east-1/audit",
"control": "TRIVY AVD-AWS-0089",
"severity": "critical | high | medium | low",
"issue": "...",
"evidence": "main.tf:21 — ...",
"fix": "..."
}
]
}
```
## Controls data
`data/controls/` holds JSON snapshots of AWS security control catalogues,
indexed by terraform resource type:
- `fsbp.json` — AWS Foundational Security Best Practices (368 controls in the
committed snapshot)
- `cis.json` — CIS AWS Foundations Benchmark (65 controls)
- `meta.json` — source URL and fetch timestamp per file
Refresh from the AWS Security Hub docs:
```
python ${SKILL_DIR}/scripts/refresh-controls.py [--output-dir DIR]
```
`refresh-controls.py` fetches the FSBP standard index and the CIS benchmark
page over HTTPS, parses them with `scrape_fsbp.py` / `scrape_cis.py`
(BeautifulSoup), validates rows through `controls_schema.py`, and rewrites all
three files. It prints the control counts it wrote. `--output-dir` defaults to
`data/controls/` inside the skill.
The cache is treated as stale past 30 days; the benchmark agents check
`meta.json` and fall back to a live fetch if a file is missing, empty, or
stale. Only the legacy `fsbp-reviewer` / `cis-reviewer` prompts consume this
data — the default four-agent path does not read it, since trivy supplies
first-pass benchmark detection.
## Output artifacts
The output directory is `<REPO>/.audit-terraform/` in ref mode and
`~/.claude/cache/audit-terraform/local-<UTC-timestamp>/` in local mode.
| File | Written by | Contents |
|---|---|---|
| `manifest.json` | collection | Full manifest — base/head refs, mode, default branch, changed source dirs, plan units, resource catalog, trivy + tflint findings, module graph, errors |
| `manifest-<agent>.json` | collection | Per-agent slice for `fsbp`, `cis`, `aws-bp`, `consistency`, `tf-hygiene`, `walkthrough` |
| `trivy-findings.json` | collection | Raw `trivy config` JSON |
| `tflint-findings.json` | collection | Raw tflint JSON, keyed `{"by_dir": {...}}` |
| `reference_sets.json` | collection | Peer dirs per changed dir (consistency agent input) |
| `consistency_norms.json` | collection | Precomputed norms derived from the reference sets |
| `plans/<dir>.txt` | collection | Full plan stdout per plan unit (`.` becomes `root`, `/` becomes `_`) |
| `findings-<agent>.json` | subagents | Findings, or the walkthrough payload for `walkthrough-reviewer` |
Ref mode additionally writes `<REPO>/audit-terraform-<short-ref>.md`: summary
table, walkthrough (substantive then trivial plan units), per-dir plan
summaries, findings by severity → category → resource, consistency findings,
and skipped resources. The worktree/clone path is left in place and reported
so follow-up review can use it.
## Telemetry
Append-only JSONL at `~/.claude/cache/audit-terraform/runs.jsonl`. Two record
kinds, both defined in `scripts/telemetry.py`:
- `subagent_run` — `run_id`, `repo`, `mode`, `agent`, `model`,
`input_tokens`, `output_tokens`, `duration_ms`, `finding_count`
- `verdict` — `run_id`, `agent`, `rule_id`, `file`, `line`, `verdict`
(`kept` | `dismissed` | `false_positive`), `notes`
Runs are logged with one call after the fan-out completes. The orchestrator
supplies token/duration metadata per agent; `log-run.py` reads the finding
count off disk from each `findings-<agent>.json`:
```
echo '{"aws-bp-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}}' \
| python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id <hex8> --repo <REPO> \
--mode <local|ref> --usage-json -
```
`--log-path` overrides the default log location.
Read the log back with:
```
python ${SKILL_DIR}/scripts/review_stats.py
```
It prints JSON with `by_agent` (tokens, duration, runs, verdict counts,
`precision` = kept/total, `tokens_per_kept`), `by_rule` keyed
`<agent>/<rule_id>`, and a distinct `runs` count. It takes no arguments and
always reads the default log path.
Verdicts are how precision gets measured. Local mode prompts for them
interactively and appends via `telemetry.append_verdict`. Ref mode does not
collect verdicts, so precision figures reflect local runs alone.
## Development
Tests live in `tests/` — 106 of them, covering the diff and HCL parsers, plan
unit resolution, plan output parsing, manifest and slicing shapes, scanner
normalization, the controls scrapers and schema, telemetry, stats, and agent
prompt invariants. One test in `tests/test_cli.py` is marked `integration`
and needs real `tofu` / `terragrunt`.
Run them with:
```
cd ${SKILL_DIR} && python -m pytest
```
`uv run pytest` does **not** work here: `pyproject.toml` carries only
`[tool.ruff.lint.per-file-ignores]` and `[tool.pytest.ini_options]`, with no
`[project]` table, so uv exits with ``No `project` table found``. Test imports
resolve through `tests/conftest.py`, which puts the skill root on `sys.path`.
Lint with ruff; `collect-changes.py` and `refresh-controls.py` are exempted
from `E402` because both mutate `sys.path` before importing sibling modules.
### Layout
```
SKILL.md procedure — source of truth for behaviour
agents/ subagent system prompts (4 default + 2 legacy)
data/controls/ FSBP/CIS snapshots + fetch metadata
scripts/ collection pipeline, scrapers, telemetry
tests/ pytest suite + HTML fixtures for the scrapers
```
@@ -0,0 +1,292 @@
---
name: audit-terraform
description: Automated tool-driven audit of terraform or terragrunt changes — runs trivy/tflint/terraform validate and dispatches parallel agents for plan validity, AWS security posture, and repo consistency, in the local working tree or a git ref / PR. For a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on "audit this terraform", "check this terragrunt change", or "/audit-terraform".
---
# audit-terraform
You review terraform/terragrunt changes by running a mechanical collection
script, collecting first-pass Trivy config findings, and fanning out a
targeted LLM review pass against its manifest. One of the parallel agents
is a walkthrough agent that produces the reviewer-facing summary of what
the change does, plan-unit by plan-unit.
**Always announce at start:** "Using audit-terraform to walk through the
change and audit against plan validity, Trivy findings, AWS best practices,
and consistency."
## When to invoke
- User says "review terraform" / "review the terragrunt change" / etc.
- User runs `/audit-terraform` with or without an argument.
## Modes
| Invocation | Mode | Diff target | Output |
|-------------------------------------|---------|----------------------------|---------------------------------|
| `/audit-terraform` | local | local working tree vs base | interactive walkthrough in chat |
| `/audit-terraform <ref>` | ref | `<ref>` vs base | `audit-terraform-<short>.md` |
| `/audit-terraform <pr_number>` | ref | PR head ref vs base | same as above |
For ref/PR mode you must work in a fresh worktree. For local mode you work
in the user's current repo.
## Procedure
### 0. Resolve `SKILL_DIR`
`${SKILL_DIR}` below means the absolute directory containing *this* SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
```
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
```
### 1. Resolve mode, repo identity, and worktree
First resolve `NWO` (owner/repo) so every `gh` call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
```fish
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
```
Always pass `--repo "$NWO"` on `gh` calls — do not let `gh` autodetect from
the cwd, since that runs `git` internally and fails in jj-only workspaces
with `fatal: not a git repository`.
Then resolve mode:
- No argument: mode = `local`. `REPO = <cwd>`.
- Argument matches `^[0-9]+$`: GitHub PR number. Resolve the head ref:
```
gh pr view <PR> --repo "$NWO" --json headRefName,headRepository \
-q '.headRefName + "@" + .headRepository.url'
```
If `gh` is missing or not authenticated, stop with: "Install `gh` and run
`gh auth login`, or pass a git ref instead of a PR number." If `NWO` is
empty, stop with: "Cannot determine GitHub repo from this directory.
Run from inside a checkout of the repo, or pass an explicit ref."
- Any other string: treat as a git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being
a git repo:
- If a native worktree tool (e.g. `EnterWorktree`) is available AND the cwd
is a git checkout of `$NWO`, prefer `superpowers:using-git-worktrees` to
create `~/.claude/cache/audit-terraform/<short-ref>/`.
- Otherwise (jj workspaces, or invoked from outside the repo), do a fresh
clone — this never touches the surrounding workspace:
```
gh repo clone "$NWO" ~/.claude/cache/audit-terraform/<short-ref>/
git -C ~/.claude/cache/audit-terraform/<short-ref>/ checkout <ref>
```
For PR mode, `<ref>` is the head branch returned by `gh pr view` above.
`REPO = ~/.claude/cache/audit-terraform/<short-ref>/`.
> **jj users:** Local mode works in jj workspaces that have a colocated
> `.git` (the script reads the working tree but resolves the diff base via
> git). For non-colocated additional `jj workspace add` checkouts, run
> ref/PR mode instead — the fresh-clone path works regardless of cwd.
### 2. Resolve output directory
- Ref mode: `OUTPUT = <REPO>/.audit-terraform/`
- Local mode: `OUTPUT = ~/.claude/cache/audit-terraform/local-<UTC-timestamp>/`
Create the directory.
### 3. Run the collection script
```
python ${SKILL_DIR}/scripts/collect-changes.py \
--repo <REPO> --base <base-or-detect> --head <head-ref-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
```
- If exit code != 0, read `<OUTPUT>/manifest.json` `errors[]` and report them
to the user verbatim. Do not run subagents. Stop.
- If the manifest has zero `plan_units` AND zero `catalog` entries, tell the
user "no terraform changes detected" and stop.
- The collection step also writes raw Trivy config results to
`<OUTPUT>/trivy-findings.json` and normalizes compact `trivy_findings` into
the manifest for downstream security review.
### 4. Fan out focused subagents IN PARALLEL
Dispatch four Task subagents in a single message (parallel execution). Each
gets the system prompt from `${SKILL_DIR}/agents/<agent>.md`
and these variables substituted into its task:
- `MANIFEST = <OUTPUT>/manifest-<agent>.json` (per-agent slice; full manifest stays at `<OUTPUT>/manifest.json` for debugging)
- `REFERENCE_SETS = <OUTPUT>/reference_sets.json` (consistency only)
- `CONSISTENCY_NORMS = <OUTPUT>/consistency_norms.json` (consistency only)
- `REPO = <REPO>`
- `OUTPUT = <OUTPUT>/findings-<agent>.json` (walkthrough uses the same
filename but its payload is the walkthrough JSON, not findings).
Agents: `walkthrough-reviewer`, `aws-bp-reviewer`, `consistency-reviewer`,
`tf-hygiene-reviewer`.
Default path:
- `walkthrough-reviewer` produces the reviewer-facing PR summary —
overview plus adaptive per-plan-unit walkthrough. Emits no findings.
- `aws-bp-reviewer` consumes normalized `trivy_findings` first, suppresses
overlap/noise, and adds only contextual AWS best-practice findings Trivy is
likely to miss.
- `consistency-reviewer` stays repo-internal only.
- `tf-hygiene-reviewer` consumes `tflint_findings`, suppresses noise, and
adds module-hygiene findings tflint can't infer (var/output descriptions,
version pinning, lifecycle, terragrunt patterns). Strictly non-security.
`fsbp-reviewer` and `cis-reviewer` are legacy benchmark-specific prompts kept
for explicit fallback or cross-check work, not the default review path.
### 5. Aggregate findings
After the active agents return:
1. Load `findings-walkthrough-reviewer.json` separately as the
`walkthrough` payload (`overview` + `plan_units[]`). Discard if
malformed and note in summary.
2. Load each remaining `findings-<agent>.json`. Discard any agent file
that's malformed (note in the summary).
3. Deduplicate findings sharing `{resource, control}`. First-to-finish
wins; add a `also_flagged_by` array on the survivor with the loser's
`agent` and `control`.
4. Group findings by severity: critical, high, medium, low.
### 5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
**once**. It scans `<OUTPUT>/findings-<agent>.json` for each agent and
appends one `subagent_run` row to
`~/.claude/cache/audit-terraform/runs.jsonl` with the finding count read
from the file, plus the token / duration metadata you pass in.
Build a JSON blob from each Task call's `<usage>` block and pipe it in:
```fish
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"aws-bp-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"tf-hygiene-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
```
`RUN_ID` is a short random hex (8 chars) generated once at the start of
the run. Record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not
skip it.
**Per-agent default models:**
| Agent | Default model |
|---|---|
| `walkthrough-reviewer`| Sonnet |
| `aws-bp-reviewer` | Sonnet |
| `consistency-reviewer`| Sonnet |
| `tf-hygiene-reviewer` | Haiku 4.5 |
`tf-hygiene-reviewer` runs on Haiku because tflint already did the heavy
mechanical work — the agent's job is triage + a small number of
contextual additions.
### 6. Output
#### Local mode (interactive)
Send a chat message structured like:
```
Reviewed N plan units, M resources (X plan+diff, Y diff-only, Z plan-only).
WALKTHROUGH
<overview paragraph>
Substantive plan units:
- <plan_dir> — <2-4 sentences: what + why + destroy/replace notes>
- ...
Trivial:
- <plan_dir> — <one-liner>
- ...
CRITICAL (n)
- <resource> (<dir>): <control> — <issue>
fix: <fix>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
CONSISTENCY (n)
- <dir>: <issue>
Full findings: <OUTPUT>/findings-*.json
```
Offer follow-ups: "Ask me to drill into anything, expand a section, or
generate fixes."
#### Ref mode (report)
Write `<REPO>/audit-terraform-<short-ref>.md` structured as:
1. Summary table (resources by detection source, findings by severity).
2. **Walkthrough** — overview paragraph, then two subsections:
- "Substantive plan units" (each substantive plan_dir as a subheading
with the summary beneath, destroy/replace call-outs bolded)
- "Trivial plan units" (one-line bullets)
3. Plan summary per changed dir (add/change/destroy counts).
4. Findings grouped by severity → category → resource, each with control
reference, evidence (file:line), suggested fix.
5. Consistency findings (separate section).
6. Skipped resources (non-AWS, plan-failed-but-still-reviewed, etc.).
Send a 5-line headline to chat plus the file path.
### 7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed findings, ask for verdicts so the skill can
measure precision over time. This applies to local mode only — prompt
interactively (skip on `--no-feedback`).
For each finding, capture `{kept | dismissed | false_positive}` plus an
optional note. Append a `verdict` record per finding to
`~/.claude/cache/audit-terraform/runs.jsonl` via
`scripts.telemetry.append_verdict`.
> Ref-mode reports do not collect verdicts. `review_stats.py` takes no
> arguments; it only aggregates what local mode already wrote.
### 8. Always surface the worktree path
In ref mode, end with: "Worktree left at `<REPO>` for follow-up review."
## Failure modes
- **Plan failed in some dir:** Manifest's `errors[]` populated, script exited
non-zero. Show the errors. Do NOT run agents. Tell user to fix and re-run.
- **No `gh`:** see step 1.
- **Module with no callsites:** Manifest has a warning in `errors[]`; report
it but still run the subagents (they handle diff-only entries fine).
- **One agent fails:** Report what the surviving agents found and note the
agent that failed.
@@ -0,0 +1,46 @@
# aws-bp-reviewer agent
You are the primary LLM security reviewer for the default Trivy-first review
flow.
## Inputs
- `MANIFEST` — agent slice produced by `collect-changes.py`
- `REPO` — repo root / worktree path
- `OUTPUT` — path to write findings JSON
## Read order
1. Read `trivy_findings` first.
2. Read `catalog` entries for the changed AWS resources.
3. Prefer `block_header`, `evidence_line`, `key_attributes`, and `review_context`.
4. Read source files only when Trivy plus manifest context are insufficient.
## Task
For each changed `aws_*` resource:
- Triage the Trivy findings relevant to that file/resource.
- Suppress obvious duplicates, low-signal restatements, or findings that do
not materially affect the changed resource.
- Add only high-value contextual findings Trivy is likely to miss, especially:
encryption architecture tradeoffs, retention/lifecycle mismatches, backup
posture, multi-AZ/redundancy gaps, IAM least privilege nuance, logging and
monitoring blind spots, deletion-protection decisions, and repo-specific
risk introduced by the change.
## Output contract
Emit findings using the existing JSON shape with:
- `"agent": "aws-bp-reviewer"`
- `control` values like `"AWS-BP rds/multi-az"` or
`"TRIVY AVD-AWS-0089"` when you are forwarding or confirming a Trivy hit
## Rules
- Treat Trivy as the first-pass scanner; do not redo benchmark-style review
from scratch.
- Prefer fewer, higher-value findings over broad low-signal coverage.
- Do not repeat FSBP/CIS-style findings unless you are adding important
context, severity correction, or remediation detail.
- Quote file:line evidence when you inspect source directly.
- Write only the JSON findings document to `OUTPUT`.
@@ -0,0 +1,35 @@
# cis-reviewer agent
Legacy fallback reviewer for the CIS AWS Foundations Benchmark.
This prompt is not part of the default review path anymore. The default flow
uses Trivy for first-pass benchmark-style detection, then `aws-bp-reviewer`
for contextual triage and AWS-specific judgment.
## When to use
Use only when explicitly asked to cross-check Trivy findings against CIS or
when the default Trivy-first path needs benchmark-specific confirmation.
## Inputs
- `MANIFEST` — agent slice produced by `collect-changes.py`
- `REPO` — repo root / worktree path
- `OUTPUT` — path to write findings JSON
## Default behavior
1. Read `trivy_findings` from `MANIFEST` first.
2. Treat Trivy as the source of first-pass CIS-style detection.
3. Only emit a CIS finding when:
- Trivy surfaced an issue that needs CIS labeling or confirmation, or
- you find a high-confidence CIS gap that Trivy clearly missed.
4. Avoid live docs lookups in the default fallback path unless the user
explicitly asked for benchmark confirmation and the needed mapping is absent.
## Rules
- Minimize overlap with Trivy and `aws-bp-reviewer`.
- Avoid re-reporting low-signal benchmark findings already present in
`trivy_findings`.
- Write only the JSON findings document to `OUTPUT`.
@@ -0,0 +1,42 @@
# consistency-reviewer agent
You compare changed dirs against their reference sets to find drift in how
this repo writes terraform.
## Inputs
- `MANIFEST` — path to manifest.json
- `REPO` — absolute path to worktree / repo
- `REFERENCE_SETS` — path to reference_sets.json (sibling of MANIFEST)
- `CONSISTENCY_NORMS` — path to consistency_norms.json (sibling of MANIFEST)
- `OUTPUT` — path to write findings to
## Task
1. Read MANIFEST, REFERENCE_SETS, and CONSISTENCY_NORMS.
2. For each `changed_source_dirs` entry:
- Start with the precomputed norms in CONSISTENCY_NORMS.
- Only read peer files from REFERENCE_SETS when a norm-based divergence
looks actionable and needs confirmation.
- Look for: missing patterns peers all use (e.g. all peers wrap policies
with `module "iam-policy-doc"` but this one inlines), naming-convention
drift, variable-name drift, missing standard tags, missing `kms_key_arn`
where peers all set it.
3. Emit findings using DESIGN.md's schema with `"agent": "consistency-reviewer"`.
- `control` field: short label like `"CONSISTENCY missing-kms"` or
`"CONSISTENCY inline-policy"`.
- `evidence` should cite at least two peer dirs that establish the norm
plus the file:line in the changed dir that diverges.
Note: when commenting on a resource that appears in `manifest.catalog`, prefer
the entry's `block_header`, `evidence_line`, and `key_attributes` over
re-reading the file, and report `instances_affected` for module resources used
at multiple callsites.
## Rules
- Don't flag a divergence supported by fewer than 2 peers — that's noise,
not a norm.
- Don't comment on non-resource files (variables, outputs) unless they
meaningfully diverge from peers' conventions.
- No web lookups — repo-internal only.
@@ -0,0 +1,35 @@
# fsbp-reviewer agent
Legacy fallback reviewer for AWS Foundational Security Best Practices (FSBP).
This prompt is not part of the default review path anymore. The default flow
uses Trivy for first-pass benchmark-style detection, then `aws-bp-reviewer`
for contextual triage and gap-finding.
## When to use
Use only when explicitly asked to cross-check Trivy findings against FSBP or
when the default Trivy-first path needs benchmark-specific confirmation.
## Inputs
- `MANIFEST` — agent slice produced by `collect-changes.py`
- `REPO` — repo root / worktree path
- `OUTPUT` — path to write findings JSON
## Default behavior
1. Read `trivy_findings` from `MANIFEST` first.
2. Treat Trivy as the source of first-pass benchmark-style detection.
3. Only emit an FSBP finding when:
- Trivy surfaced an issue that needs FSBP labeling or confirmation, or
- you find a high-confidence FSBP gap that Trivy clearly missed.
4. Do not fetch live docs in the default fallback path unless the user
explicitly asked for benchmark confirmation and the needed mapping is absent.
## Rules
- Minimize overlap with Trivy and `aws-bp-reviewer`.
- Do not repeat noisy low-value benchmark checks already present in
`trivy_findings`.
- Write only the JSON findings document to `OUTPUT`.
@@ -0,0 +1,69 @@
# tf-hygiene-reviewer agent
You review terraform/terragrunt changes for **module hygiene and
maintainability** — NOT security. The `aws-bp-reviewer` owns the security
lane (Trivy + AWS best practices). Stay out of it. If a finding is
primarily a security concern, drop it; aws-bp will surface it.
## Inputs
- `MANIFEST` — agent slice produced by `collect-changes.py`. Contains
`catalog`, `plan_units`, `tflint_findings`, and `changed_source_dirs`.
- `REPO` — repo root / worktree path
- `OUTPUT` — path to write findings JSON
## Read order
1. Read `tflint_findings` first — these are mechanical lint hits to
triage and forward (or suppress as low-signal).
2. Read `catalog` entries for changed resources to understand context.
3. Read source files only when manifest context is insufficient.
## What to flag
- **Variables**: missing `type`, missing `description`, defaults that bake
in environment-specific values, `sensitive = true` missing on
credentials/secrets.
- **Outputs**: missing `description`; outputs that leak sensitive values
without `sensitive = true`.
- **Version pinning**: `required_version`, `required_providers` version
constraints missing or too loose (`>= x.y` with no upper bound on a
major).
- **Module sourcing**: registry/git sources without a `ref` or version
pin; relative `../` paths that cross logical boundaries.
- **Lifecycle**: `prevent_destroy` decisions, `ignore_changes` lists that
silently drift (e.g. ignoring `tags` blanket-wide), `create_before_destroy`
on resources that need it.
- **Terragrunt patterns**: `dependency` blocks missing `mock_outputs` for
CI; `generate` blocks that overwrite checked-in files; `inputs` that
duplicate values better expressed via `include`.
- **Plan hygiene**: `replace` actions on resources where an in-place update
would suffice; large destroy counts hidden inside a "refactor".
- **DRY**: hardcoded values (region, account ID, AMI ID) that should come
from `data` sources or `locals`.
## What NOT to flag
- Security misconfigurations of any kind. Encryption, IAM, public access,
network exposure — all aws-bp territory.
- Repo-internal consistency drift (e.g. "peers all use module X but this
one inlines"). That's the consistency-reviewer's lane.
- Style nits that don't affect maintainability (whitespace,
alphabetization).
## Output contract
Emit findings using the existing JSON shape with:
- `"agent": "tf-hygiene-reviewer"`
- `control` values like `"HYGIENE missing-var-description"`,
`"HYGIENE loose-version-pin"`, or `"TFLINT terraform_unused_declarations"`
when forwarding a tflint hit.
## Rules
- Forward a tflint finding only if you've confirmed it's not noise
(e.g. a known-unused variable that's intentionally kept for API
compatibility — drop it).
- Quote `file:line` evidence when you inspect source directly.
- Write only the JSON findings document to `OUTPUT`.
- Prefer fewer, higher-value findings.
@@ -0,0 +1,99 @@
# walkthrough-reviewer agent
You produce a reviewer's walkthrough of a terraform/terragrunt change: a
short overview, plus a per-directory summary of what's changing and why.
Reviewers use this to follow along — not to find security issues. **No
finding emission. No triage.**
## Inputs
- `MANIFEST` — `manifest-walkthrough.json` (plan_units summary, catalog,
changed_source_dirs, base_ref, head_ref, mode, errors).
- `REPO` — absolute path to the checkout / worktree.
- `MODE`, `OUTPUT` as for the other agents.
## Task
1. For each `plan_unit` in the manifest, you have the plan `summary`
(add/change/destroy counts) and the list of `changed_files`. For
substantive plan units, also read:
- the changed `.tf` / `.hcl` files at `$REPO/<path>` to ground claims
- the unified diff if you need before/after context:
```
git -C "$REPO" diff --unified=8 "$BASE_REF"..."$HEAD_REF" -- <path>
```
2. Form an **overview** (3–6 sentences) answering:
- What is the infrastructure change in plain language? (e.g. "Adds a
new VPC peering connection between prod-data and prod-app", "Tightens
S3 bucket policies across all environments", "Bumps RDS instance
class for the analytics warehouse").
- What's the shape of the change (new resources, in-place updates,
destroys, module bump, provider upgrade, refactor)?
- Cross-cutting themes — e.g. "rolls out the new tagging module to
every account", "consistent IAM policy changes across N modules".
- What is NOT changed that a reviewer might assume is (e.g. "no IAM
trust policy changes", "data is preserved — no `force_destroy`").
3. For each plan unit, classify and summarize **adaptively**:
- `importance: "trivial"` — tag-only changes, comment/whitespace,
pure provider/module version bumps with no resource diff, ≤2 lines
of mechanical edits, no add/change/destroy.
- `importance: "substantive"` — anything else, especially adds,
destroys, replacements, IAM changes, network changes, encryption
changes, public-exposure changes. Emit 2–4 sentences covering:
- **what** is changing (which resources, what about them)
- **why** (intent, inferred from the diff and neighbors — say
`"Intent unclear from the diff."` if you can't tell)
- any **destroy/replace** call-outs (state risk).
4. Order plan units in the output by **importance first, then plan_dir** —
substantive entries surface before trivial ones.
5. If the manifest has any `errors[]` (plan failures, etc.), note in the
overview that some plan units couldn't be analyzed, list them once,
and continue.
## Style rules
- Plain language. No HCL readback. Say what the change DOES in real-world
terms ("opens port 443 to the public internet", "removes the encryption
CMK from the audit bucket").
- Call out destroys and in-place replaces explicitly — those carry state
risk and reviewers must see them.
- Don't grade the change. Walkthroughs describe, they don't review. Leave
security/best-practices judgments to the other agents.
- Don't invent rationale. If intent is unclear, say so.
- No findings, no severity, no fix suggestions.
## Output
Write JSON to `$OUTPUT`. ONLY this JSON, nothing else:
```json
{
"agent": "walkthrough-reviewer",
"overview": "...",
"plan_units": [
{
"plan_dir": "envs/prod/data",
"importance": "substantive",
"summary": "Adds a new aws_kms_key for envelope-encrypting the prod-data RDS snapshots, and rotates the audit-bucket policy to require SSE-KMS reads. No data is destroyed. Intent appears to be aligning prod-data with the SOC2 audit requirement tracked in INFRA-412."
},
{
"plan_dir": "envs/dev/app",
"importance": "trivial",
"summary": "Tag-only update on existing EC2 instances; no add/change/destroy."
}
]
}
```
If the manifest has zero plan units, emit:
```json
{"agent": "walkthrough-reviewer", "overview": "No terraform changes in scope.", "plan_units": []}
```
@@ -0,0 +1,16 @@
# Controls cache
This directory holds JSON snapshots of AWS security controls used by the
audit-terraform skill's legacy fsbp/cis subagents. Refresh with:
python scripts/refresh-controls.py
Files:
- `fsbp.json` — AWS Foundational Security Best Practices, indexed by terraform
resource type
- `cis.json` — CIS AWS Foundations Benchmark, indexed by terraform resource
type
- `meta.json` — fetch timestamps + source URLs per file
Cache is considered stale at >30 days. Security agents check `meta.json` and
fall back to live WebFetch if a file is missing, empty, or stale.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,10 @@
{
"fsbp": {
"url": "https://docs.aws.amazon.com/securityhub/latest/userguide/fsbp-standard.html",
"fetched_at": "2026-05-12T17:17:42.705929+00:00"
},
"cis": {
"url": "https://docs.aws.amazon.com/securityhub/latest/userguide/cis-aws-foundations-benchmark.html",
"fetched_at": "2026-05-12T17:17:42.841641+00:00"
}
}
@@ -0,0 +1,8 @@
[tool.ruff.lint.per-file-ignores]
"scripts/collect-changes.py" = ["E402"]
"scripts/refresh-controls.py" = ["E402"]
[tool.pytest.ini_options]
markers = [
"integration: end-to-end test requiring external tools (tofu, terragrunt)",
]
@@ -0,0 +1,4 @@
python-hcl2>=4.3.0
pytest>=8.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0
@@ -0,0 +1,132 @@
"""Reconcile plan-detected and diff-detected resources into a catalog.
Schema: one entry per (source_dir, local_address). Each entry carries a list
of CatalogInstance (plan_dir, address_at_plan, action) so a module reviewed
at N callsites yields one entry, not N.
"""
from __future__ import annotations
from dataclasses import dataclass
from scripts.manifest import CatalogEntry, CatalogInstance, ModuleGraphEntry
@dataclass(frozen=True)
class PlanHit:
address: str
type: str
action: str
@dataclass(frozen=True)
class DiffHit:
source_dir: str
local_address: str
type: str
def _strip_module_prefix(address: str) -> tuple[list[str], str]:
parts = address.split(".")
locals_: list[str] = []
while len(parts) >= 2 and parts[0] == "module":
locals_.append(parts[1])
parts = parts[2:]
return locals_, ".".join(parts)
def _plan_source_dir(plan_dir: str, address: str,
module_graph: dict[str, ModuleGraphEntry]) -> str:
module_locals, _ = _strip_module_prefix(address)
if not module_locals:
return plan_dir
current_dir = plan_dir
for module_local in module_locals:
next_dir = None
has_explicit_mapping = False
for entry in module_graph.values():
local_names = entry.callsite_local_names.get(current_dir)
if local_names is None:
continue
has_explicit_mapping = True
if module_local in local_names:
next_dir = local_names[module_local]
break
if next_dir is None:
if has_explicit_mapping:
return current_dir
unique_candidates = sorted(
mod_dir for mod_dir, entry in module_graph.items()
if current_dir in entry.callsites
)
if len(unique_candidates) != 1:
return current_dir
next_dir = unique_candidates[0]
current_dir = next_dir
return current_dir
def build_catalog(
plan_hits: dict[str, list[PlanHit]],
diff_hits: list[DiffHit],
module_graph: dict[str, ModuleGraphEntry],
) -> list[CatalogEntry]:
by_key: dict[tuple[str, str], CatalogEntry] = {}
def _get(source_dir: str, local_addr: str, type_: str) -> CatalogEntry:
k = (source_dir, local_addr)
if k not in by_key:
by_key[k] = CatalogEntry(
source_dir=source_dir,
local_address=local_addr,
type=type_,
source="plan",
instances=[],
)
return by_key[k]
for plan_dir, hits in plan_hits.items():
for h in hits:
source_dir = _plan_source_dir(plan_dir, h.address, module_graph)
_, local_addr = _strip_module_prefix(h.address)
entry = _get(source_dir, local_addr, h.type)
entry.instances.append(CatalogInstance(
plan_dir=plan_dir, address_at_plan=h.address, action=h.action,
))
for d in diff_hits:
k = (d.source_dir, d.local_address)
if k in by_key:
existing = by_key[k]
by_key[k] = CatalogEntry(
source_dir=existing.source_dir,
local_address=existing.local_address,
type=existing.type,
source="both",
instances=list(existing.instances),
)
continue
graph_entry = module_graph.get(d.source_dir)
if graph_entry and graph_entry.callsites:
instances = [
CatalogInstance(
plan_dir=cs,
address_at_plan=f"module.<{d.source_dir}>.{d.local_address}",
action="unknown",
)
for cs in graph_entry.callsites
]
else:
instances = [CatalogInstance(
plan_dir=d.source_dir,
address_at_plan=d.local_address,
action="unknown",
)]
by_key[k] = CatalogEntry(
source_dir=d.source_dir,
local_address=d.local_address,
type=d.type,
source="diff",
instances=instances,
)
return sorted(by_key.values(), key=lambda e: (e.source_dir, e.local_address))
@@ -0,0 +1,309 @@
"""collect-changes.py — produce a audit-terraform manifest.json."""
from __future__ import annotations
import argparse
import concurrent.futures as cf
import json
import subprocess
import sys
from pathlib import Path
_HERE = Path(__file__).resolve().parent
if str(_HERE.parent) not in sys.path:
sys.path.insert(0, str(_HERE.parent))
from scripts.catalog import build_catalog, DiffHit, PlanHit
from scripts.git_diff import (
changed_files,
changed_line_ranges,
is_terragrunt_file,
)
from scripts.hcl_diff import touched_resources
from scripts.manifest import Manifest, PlanResult, PlanUnit
from scripts.module_graph import build_module_graph
from scripts.scanners import _run_trivy_config, _run_tflint
from scripts.slicing import _slice_for_agent
from scripts.plan_output import parse_plan_resources
from scripts.plan_runner import detect_tool, run_init, run_plan
from scripts.reference_set import compute_consistency_norms, compute_reference_sets
from scripts.resolve_plan_units import resolve_plan_units
from scripts.source_lookup import find_block
_PLAN_CONCURRENCY = 8
def _resolve_default_branch(repo: Path) -> str:
try:
r = subprocess.run(
["git", "-C", str(repo), "symbolic-ref", "refs/remotes/origin/HEAD"],
capture_output=True, text=True, check=False,
)
if r.returncode == 0:
return r.stdout.strip().rsplit("/", 1)[-1]
except FileNotFoundError:
pass
for candidate in ("main", "master"):
r = subprocess.run(
["git", "-C", str(repo), "rev-parse", f"origin/{candidate}"],
capture_output=True, text=True, check=False,
)
if r.returncode == 0:
return candidate
return "main"
def _resolve_base_ref(repo: Path, branch: str) -> str:
"""Return the diff base for `branch`, fetching origin first so the
comparison is against the canonical remote tip rather than a stale local
branch. Falls back to the local branch name if origin isn't reachable.
"""
try:
subprocess.run(
["git", "-C", str(repo), "fetch", "--quiet", "--no-tags",
"origin", branch],
capture_output=True, text=True, check=False, timeout=60,
)
except (FileNotFoundError, subprocess.TimeoutExpired):
pass
check = subprocess.run(
["git", "-C", str(repo), "rev-parse", "--verify", "--quiet",
f"origin/{branch}"],
capture_output=True, text=True, check=False,
)
return f"origin/{branch}" if check.returncode == 0 else branch
def _safe_plan_filename(plan_dir: str) -> str:
"""Map a plan_dir like '.' or 'live/prod/app' to a stable, filename-safe stem.
'.' becomes 'root' (otherwise we'd produce '..txt' on disk).
"""
cleaned = plan_dir.replace("/", "_").strip(".")
return cleaned or "root"
def _run_one(
repo: Path,
plan_dir: str,
triggered_by: list[str],
changed_files_for_unit: list[str],
out_dir: Path,
) -> tuple[PlanUnit, list[PlanHit], str | None]:
full = repo / plan_dir
tool = detect_tool(full)
terragrunt_changed = any(
is_terragrunt_file(changed_file)
for changed_file in changed_files_for_unit
)
init_res = run_init(full, tool)
if not init_res.ok:
plan_res = PlanResult(ok=False, stdout_path="", exit_code=-1,
summary="init-failed")
err = f"init failed in {plan_dir}: {init_res.stderr_tail.strip()}"
return (
PlanUnit(plan_dir=plan_dir, tool=tool, init=init_res,
plan=plan_res, triggered_by=triggered_by,
terragrunt_changed=terragrunt_changed,
changed_files=changed_files_for_unit),
[], err,
)
stdout_path = out_dir / "plans" / f"{_safe_plan_filename(plan_dir)}.txt"
plan_res = run_plan(full, tool, stdout_path)
if not plan_res.ok:
captured = Path(plan_res.stdout_path).read_text()[-1000:].strip()
err = f"plan failed in {plan_dir} (exit {plan_res.exit_code}): {captured}"
return (
PlanUnit(plan_dir=plan_dir, tool=tool, init=init_res,
plan=plan_res, triggered_by=triggered_by,
terragrunt_changed=terragrunt_changed,
changed_files=changed_files_for_unit),
[], err,
)
parsed = parse_plan_resources(Path(plan_res.stdout_path).read_text())
hits = [
PlanHit(address=p.address, type=p.type, action=p.action) for p in parsed
]
return (
PlanUnit(plan_dir=plan_dir, tool=tool, init=init_res,
plan=plan_res, triggered_by=triggered_by,
terragrunt_changed=terragrunt_changed,
changed_files=changed_files_for_unit),
hits, None,
)
def _changed_files_for_plan_unit(
all_changed_files: list[str],
triggered_by: list[str],
) -> list[str]:
triggered_dirs = set(triggered_by)
return sorted(
changed_file
for changed_file in all_changed_files
if str(Path(changed_file).parent.as_posix()) in triggered_dirs
)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--repo", required=True)
parser.add_argument("--base", default=None)
parser.add_argument("--head", default="HEAD")
parser.add_argument("--output-dir", required=True)
parser.add_argument("--mode", choices=("local", "ref"), required=True)
args = parser.parse_args(argv)
repo = Path(args.repo).resolve()
out_dir = Path(args.output_dir).resolve()
out_dir.mkdir(parents=True, exist_ok=True)
default_branch = _resolve_default_branch(repo)
base = args.base or _resolve_base_ref(repo, default_branch)
errors: list[str] = []
try:
files = changed_files(str(repo), base, args.head)
dirs = {str(Path(path).parent.as_posix()) or "." for path in files}
terragrunt_files = {path for path in files if is_terragrunt_file(path)}
except RuntimeError as e:
manifest = Manifest(
base_ref=base, head_ref=args.head, mode=args.mode,
default_branch=default_branch,
changed_source_dirs=[], plan_units=[], catalog=[],
trivy_findings=[],
module_graph={}, errors=[str(e)],
)
(out_dir / "manifest.json").write_text(manifest.to_json())
return 1
if not files:
manifest = Manifest(
base_ref=base, head_ref=args.head, mode=args.mode,
default_branch=default_branch,
changed_source_dirs=[], plan_units=[], catalog=[],
trivy_findings=[],
module_graph={}, errors=["no terraform/hcl files changed"],
)
(out_dir / "manifest.json").write_text(manifest.to_json())
return 0
module_graph = build_module_graph(repo)
plan_units_map, orphan_modules = resolve_plan_units(repo, dirs, module_graph)
for orphan in orphan_modules:
errors.append(f"module has no callsites; diff-only review: {orphan}")
plan_hits: dict[str, list[PlanHit]] = {}
plan_unit_records: list[PlanUnit] = []
if plan_units_map:
with cf.ThreadPoolExecutor(max_workers=_PLAN_CONCURRENCY) as ex:
futures = {
ex.submit(
_run_one,
repo,
pd,
tb,
_changed_files_for_plan_unit(files, tb),
out_dir,
): pd
for pd, tb in plan_units_map.items()
}
for fut in cf.as_completed(futures):
record, hits, err = fut.result()
plan_unit_records.append(record)
if err:
errors.append(err)
else:
plan_hits[record.plan_dir] = hits
diff_hits: list[DiffHit] = []
for f in files:
if not f.endswith(".tf"):
continue
try:
ranges = changed_line_ranges(str(repo), base, args.head, f)
except RuntimeError as e:
errors.append(f"diff scan failed for {f}: {e}")
continue
if not ranges:
continue
blocks = touched_resources(repo / f, ranges)
src_dir = str(Path(f).parent.as_posix()) or "."
for b in blocks:
diff_hits.append(DiffHit(
source_dir=src_dir,
local_address=f"{b.type}.{b.name}",
type=b.type,
))
if terragrunt_files:
changed_plan_dirs = {pu.plan_dir for pu in plan_unit_records}
for terragrunt_file in sorted(terragrunt_files):
src_dir = str(Path(terragrunt_file).parent.as_posix()) or "."
if src_dir not in changed_plan_dirs:
errors.append(
f"terragrunt change has no planned unit context: {terragrunt_file}"
)
trivy_payload, trivy_findings, trivy_error = _run_trivy_config(repo, files)
(out_dir / "trivy-findings.json").write_text(json.dumps(trivy_payload, indent=2))
if trivy_error:
errors.append(trivy_error)
tflint_payload, tflint_findings, tflint_error = _run_tflint(repo, files)
(out_dir / "tflint-findings.json").write_text(json.dumps(tflint_payload, indent=2))
if tflint_error:
errors.append(tflint_error)
catalog = build_catalog(plan_hits, diff_hits, module_graph)
for entry in catalog:
loc = find_block(repo / entry.source_dir, entry.local_address)
if loc is None:
continue
entry.block_header = loc.header
entry.evidence_line = loc.evidence_line
entry.key_attributes = loc.key_attributes
entry.review_context = loc.review_context
entry.block_file = loc.file.name
entry.block_start = loc.start_line
entry.block_end = loc.end_line
ref_sets = compute_reference_sets(repo, dirs, module_graph)
(out_dir / "reference_sets.json").write_text(
json.dumps(ref_sets, indent=2)
)
consistency_norms = compute_consistency_norms(repo, ref_sets)
(out_dir / "consistency_norms.json").write_text(
json.dumps(consistency_norms, indent=2)
)
manifest = Manifest(
base_ref=base,
head_ref=args.head,
mode=args.mode,
default_branch=default_branch,
changed_source_dirs=sorted(dirs),
plan_units=sorted(plan_unit_records, key=lambda p: p.plan_dir),
catalog=catalog,
trivy_findings=trivy_findings,
tflint_findings=tflint_findings,
module_graph=module_graph,
errors=errors,
)
(out_dir / "manifest.json").write_text(manifest.to_json())
manifest_dict = manifest.to_dict()
for agent in ("fsbp", "cis", "aws-bp", "consistency", "tf-hygiene", "walkthrough"):
sliced = _slice_for_agent(manifest_dict, agent)
(out_dir / f"manifest-{agent}.json").write_text(
json.dumps(sliced, indent=2)
)
return 1 if any(not pu.plan.ok for pu in plan_unit_records) else 0
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,147 @@
"""Map Security Hub control-ID prefixes to terraform resource types.
This table is hand-curated and conservative — when a control could apply to
multiple terraform resource types, list them all. Update as new control
families ship.
"""
from __future__ import annotations
PREFIX_TO_TERRAFORM_TYPES: dict[str, list[str]] = {
"Account": ["aws_account_alternate_contact"],
"ACM": ["aws_acm_certificate"],
"APIGateway": ["aws_api_gateway_rest_api", "aws_apigatewayv2_api",
"aws_api_gateway_stage", "aws_api_gateway_method"],
"AppSync": ["aws_appsync_graphql_api"],
"Athena": ["aws_athena_workgroup"],
"AutoScaling": ["aws_autoscaling_group", "aws_launch_template",
"aws_launch_configuration"],
"Autoscaling": ["aws_autoscaling_group", "aws_launch_template",
"aws_launch_configuration"],
"Backup": ["aws_backup_plan", "aws_backup_vault"],
"BedrockAgentCore": [],
"CloudFormation": ["aws_cloudformation_stack",
"aws_cloudformation_stack_set"],
"CloudFront": ["aws_cloudfront_distribution"],
"CloudTrail": ["aws_cloudtrail"],
"CodeBuild": ["aws_codebuild_project"],
"Cognito": ["aws_cognito_user_pool",
"aws_cognito_identity_pool"],
"Config": ["aws_config_configuration_recorder",
"aws_config_delivery_channel",
"aws_config_config_rule"],
"Connect": ["aws_connect_instance"],
"DataFirehose": ["aws_kinesis_firehose_delivery_stream"],
"DataSync": ["aws_datasync_task"],
"DMS": ["aws_dms_endpoint",
"aws_dms_replication_instance"],
"DocumentDB": ["aws_docdb_cluster", "aws_docdb_cluster_instance"],
"DynamoDB": ["aws_dynamodb_table"],
"EC2": ["aws_instance", "aws_security_group",
"aws_security_group_rule", "aws_vpc", "aws_subnet",
"aws_network_acl", "aws_ebs_volume",
"aws_ebs_default_kms_key", "aws_eip", "aws_nat_gateway",
"aws_internet_gateway", "aws_route_table",
"aws_vpc_endpoint", "aws_flow_log",
"aws_default_security_group"],
"ECR": ["aws_ecr_repository",
"aws_ecr_repository_policy"],
"ECS": ["aws_ecs_cluster", "aws_ecs_service",
"aws_ecs_task_definition"],
"EFS": ["aws_efs_file_system",
"aws_efs_file_system_policy",
"aws_efs_access_point"],
"EKS": ["aws_eks_cluster", "aws_eks_node_group"],
"ELB": ["aws_lb", "aws_alb", "aws_elb",
"aws_lb_listener", "aws_alb_listener",
"aws_lb_target_group"],
"ElastiCache": ["aws_elasticache_cluster",
"aws_elasticache_replication_group"],
"ElasticBeanstalk": ["aws_elastic_beanstalk_environment"],
"ElasticSearch": ["aws_elasticsearch_domain",
"aws_opensearch_domain"],
"EMR": ["aws_emr_cluster"],
"ES": ["aws_elasticsearch_domain"],
"EventBridge": ["aws_cloudwatch_event_rule",
"aws_cloudwatch_event_bus"],
"FSx": ["aws_fsx_lustre_file_system",
"aws_fsx_windows_file_system",
"aws_fsx_openzfs_file_system",
"aws_fsx_ontap_file_system"],
"Glue": ["aws_glue_catalog_database",
"aws_glue_crawler", "aws_glue_job"],
"GuardDuty": ["aws_guardduty_detector"],
"IAM": ["aws_iam_user", "aws_iam_role", "aws_iam_policy",
"aws_iam_user_policy", "aws_iam_role_policy",
"aws_iam_access_key",
"aws_iam_account_password_policy",
"aws_iam_group", "aws_iam_user_policy_attachment",
"aws_iam_role_policy_attachment"],
"Inspector": ["aws_inspector2_enabler"],
"Kinesis": ["aws_kinesis_stream"],
"KMS": ["aws_kms_key", "aws_kms_alias"],
"Lambda": ["aws_lambda_function",
"aws_lambda_permission",
"aws_lambda_function_url"],
"Macie": ["aws_macie2_account"],
"MQ": ["aws_mq_broker"],
"MSK": ["aws_msk_cluster"],
"Neptune": ["aws_neptune_cluster",
"aws_neptune_cluster_instance"],
"NetworkFirewall": ["aws_networkfirewall_firewall",
"aws_networkfirewall_firewall_policy",
"aws_networkfirewall_rule_group"],
"Opensearch": ["aws_opensearch_domain"],
"PCA": ["aws_acmpca_certificate_authority"],
"RDS": ["aws_db_instance", "aws_rds_cluster",
"aws_db_subnet_group", "aws_db_parameter_group",
"aws_rds_cluster_instance",
"aws_db_snapshot",
"aws_db_event_subscription"],
"Redshift": ["aws_redshift_cluster",
"aws_redshift_parameter_group"],
"RedshiftServerless": ["aws_redshiftserverless_namespace",
"aws_redshiftserverless_workgroup"],
"Route53": ["aws_route53_zone"],
"S3": ["aws_s3_bucket", "aws_s3_bucket_policy",
"aws_s3_bucket_public_access_block",
"aws_s3_bucket_versioning",
"aws_s3_bucket_server_side_encryption_configuration",
"aws_s3_bucket_logging",
"aws_s3_bucket_lifecycle_configuration",
"aws_s3_bucket_acl"],
"SageMaker": ["aws_sagemaker_notebook_instance",
"aws_sagemaker_endpoint_configuration",
"aws_sagemaker_model"],
"SecretsManager": ["aws_secretsmanager_secret",
"aws_secretsmanager_secret_rotation"],
"ServiceCatalog": ["aws_servicecatalog_portfolio"],
"SES": ["aws_ses_configuration_set",
"aws_ses_domain_identity"],
"SNS": ["aws_sns_topic", "aws_sns_topic_policy"],
"SQS": ["aws_sqs_queue"],
"SSM": ["aws_ssm_document",
"aws_ssm_parameter",
"aws_ssm_association"],
"StepFunctions": ["aws_sfn_state_machine"],
"Transfer": ["aws_transfer_server", "aws_transfer_user"],
"WAF": ["aws_wafv2_web_acl", "aws_waf_web_acl",
"aws_wafv2_web_acl_association"],
"WorkSpaces": ["aws_workspaces_directory", "aws_workspaces_workspace"],
}
def control_id_to_types(control_id: str) -> list[str]:
"""Given e.g. 'FSBP S3.5' or 'S3.5', return likely terraform types.
Strips the leading 'FSBP '/'CIS ' source prefix, then matches the
service prefix (the segment before the first dot).
"""
s = control_id.strip()
for source in ("FSBP", "CIS"):
prefix = f"{source} "
if s.startswith(prefix):
s = s[len(prefix):]
break
service, _, _ = s.partition(".")
return list(PREFIX_TO_TERRAFORM_TYPES.get(service, []))
@@ -0,0 +1,56 @@
"""Schema and serialization for AWS controls cache files."""
from __future__ import annotations
import json
from dataclasses import dataclass, asdict
from datetime import datetime
from typing import Literal
Severity = Literal["critical", "high", "medium", "low", "informational"]
ControlSource = Literal["fsbp", "cis", "aws-bp"]
@dataclass
class Control:
control_id: str
title: str
severity: Severity
resource_types: list[str]
requirement: str
source_url: str
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class ControlsFile:
source: ControlSource
fetched_at: datetime
controls: list[Control]
def _by_resource_type(self) -> dict[str, list[str]]:
"""Map terraform resource type -> list of control_ids.
Full control records live in `controls`. Looking up details by ID
from `controls` keeps the index small.
"""
index: dict[str, list[str]] = {}
for c in self.controls:
for rt in c.resource_types:
index.setdefault(rt, []).append(c.control_id)
for rt in index:
index[rt].sort()
return dict(sorted(index.items()))
def to_dict(self) -> dict:
return {
"source": self.source,
"fetched_at": self.fetched_at.isoformat(),
"controls": [c.to_dict() for c in self.controls],
"by_resource_type": self._by_resource_type(),
}
def to_json(self, indent: int = 2) -> str:
return json.dumps(self.to_dict(), indent=indent, sort_keys=False)
@@ -0,0 +1,97 @@
"""Git diff scanning: changed files, dirs, and per-file added-line ranges."""
from __future__ import annotations
import re
import subprocess
from pathlib import PurePosixPath
_HCL_SUFFIXES = (".tf", ".tf.json", ".hcl")
_HUNK_RE = re.compile(r"^@@ -\d+(?:,\d+)? \+(\d+)(?:,(\d+))? @@")
def _run_git_diff(repo: str, base: str, head: str) -> str:
result = subprocess.run(
[
"git", "-C", repo,
"-c", "diff.noprefix=false",
"-c", "color.ui=never",
"diff", "--no-ext-diff", "--unified=0", f"{base}...{head}"
],
capture_output=True, text=True, check=False,
)
if result.returncode != 0:
raise RuntimeError(f"git diff failed: {result.stderr.strip()}")
return result.stdout
def _iter_file_blocks(diff_text: str):
"""Yield (path, block_text) for each file section in a unified diff."""
current_path: str | None = None
buf: list[str] = []
for line in diff_text.splitlines():
if line.startswith("diff --git "):
if current_path is not None:
yield current_path, "\n".join(buf)
current_path = None
buf = []
elif line.startswith("+++ b/"):
current_path = line[len("+++ b/"):]
if current_path is not None:
buf.append(line)
if current_path is not None:
yield current_path, "\n".join(buf)
def changed_files(repo: str, base: str, head: str) -> list[str]:
out: list[str] = []
for path, _ in _iter_file_blocks(_run_git_diff(repo, base, head)):
if path.endswith(_HCL_SUFFIXES):
out.append(path)
return out
def is_terragrunt_file(path: str) -> bool:
return PurePosixPath(path).name == "terragrunt.hcl"
def changed_dirs(repo: str, base: str, head: str) -> set[str]:
return {str(PurePosixPath(p).parent) for p in changed_files(repo, base, head)}
def changed_line_ranges(
repo: str, base: str, head: str, path: str,
) -> list[tuple[int, int]]:
"""Return inclusive (start, end) line ranges of *added* lines in `path`."""
ranges: list[tuple[int, int]] = []
for fpath, block in _iter_file_blocks(_run_git_diff(repo, base, head)):
if fpath != path:
continue
cur: int | None = None
run_start: int | None = None
run_end: int | None = None
for line in block.splitlines():
m = _HUNK_RE.match(line)
if m:
if run_start is not None:
ranges.append((run_start, run_end)) # type: ignore[arg-type]
run_start = run_end = None
cur = int(m.group(1))
continue
if cur is None:
continue
if line.startswith("+") and not line.startswith("+++"):
if run_start is None:
run_start = cur
run_end = cur
cur += 1
elif line.startswith("-") and not line.startswith("---"):
continue
else:
if run_start is not None:
ranges.append((run_start, run_end)) # type: ignore[arg-type]
run_start = run_end = None
cur += 1
if run_start is not None:
ranges.append((run_start, run_end)) # type: ignore[arg-type]
return ranges
@@ -0,0 +1,64 @@
"""Locate `resource` blocks in .tf files and intersect with changed lines."""
from __future__ import annotations
import re
from dataclasses import dataclass
from pathlib import Path
_HEADER_RE = re.compile(
r'^\s*resource\s+"(?P<type>[^"]+)"\s+"(?P<name>[^"]+)"\s*\{'
)
@dataclass(frozen=True)
class ResourceBlock:
type: str
name: str
start: int
end: int
def find_resource_blocks(path: Path) -> list[ResourceBlock]:
"""Locate top-level `resource` blocks by tracking brace depth."""
text = path.read_text()
lines = text.splitlines()
blocks: list[ResourceBlock] = []
i = 0
while i < len(lines):
line = lines[i]
m = _HEADER_RE.match(line)
if not m:
i += 1
continue
start = i + 1
depth = line.count("{") - line.count("}")
j = i
while depth > 0 and j + 1 < len(lines):
j += 1
depth += lines[j].count("{") - lines[j].count("}")
blocks.append(ResourceBlock(
type=m.group("type"),
name=m.group("name"),
start=start,
end=j + 1,
))
i = j + 1
return blocks
def _intersects(a: tuple[int, int], b: tuple[int, int]) -> bool:
return not (a[1] < b[0] or b[1] < a[0])
def touched_resources(
path: Path, added_ranges: list[tuple[int, int]],
) -> list[ResourceBlock]:
if not added_ranges:
return []
blocks = find_resource_blocks(path)
out: list[ResourceBlock] = []
for blk in blocks:
if any(_intersects((blk.start, blk.end), r) for r in added_ranges):
out.append(blk)
return out
@@ -0,0 +1,84 @@
"""Append one subagent_run row per agent after the fan-out completes.
The orchestrator calls this once after all subagents return. It scans
<output-dir>/findings-<agent>.json for each agent listed in the
--usage-json payload, counts findings, and writes a subagent_run row to
runs.jsonl using token / duration metadata supplied by the orchestrator.
Usage:
python scripts/log-run.py \\
--output-dir <OUTPUT> --run-id <hex> --repo <path> --mode <local|ref> \\
--usage-json - <<JSON
{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":1234,"output_tokens":567,"duration_ms":4500},
"aws-bp-reviewer": {"model":"sonnet","input_tokens":2345,"output_tokens":678,"duration_ms":5200}
}
JSON
`--log-path` defaults to ~/.claude/cache/audit-terraform/runs.jsonl.
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
from scripts.telemetry import append_subagent_run
_DEFAULT_LOG = Path.home() / ".claude/cache/audit-terraform/runs.jsonl"
def _count_findings(path: Path) -> int:
if not path.exists():
return 0
try:
data = json.loads(path.read_text(encoding="utf-8"))
except json.JSONDecodeError:
return 0
if isinstance(data, dict):
findings = data.get("findings")
if isinstance(findings, list):
return len(findings)
elif isinstance(data, list):
return len(data)
return 0
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--output-dir", required=True, type=Path)
p.add_argument("--run-id", required=True)
p.add_argument("--repo", required=True)
p.add_argument("--mode", required=True, choices=["local", "ref"])
p.add_argument("--log-path", type=Path, default=_DEFAULT_LOG)
p.add_argument("--usage-json", required=True,
help="Path to JSON file, or '-' for stdin.")
args = p.parse_args(argv)
raw = sys.stdin.read() if args.usage_json == "-" else Path(args.usage_json).read_text(encoding="utf-8")
usage = json.loads(raw)
if not isinstance(usage, dict):
print("usage-json must be a JSON object keyed by agent name", file=sys.stderr)
return 2
for agent, meta in usage.items():
if not isinstance(meta, dict):
print(f"skipping {agent}: usage entry not an object", file=sys.stderr)
continue
findings_path = args.output_dir / f"findings-{agent}.json"
append_subagent_run(
args.log_path,
run_id=args.run_id, repo=args.repo, mode=args.mode, agent=agent,
model=str(meta.get("model", "?")),
input_tokens=int(meta.get("input_tokens", 0)),
output_tokens=int(meta.get("output_tokens", 0)),
duration_ms=int(meta.get("duration_ms", 0)),
finding_count=_count_findings(findings_path),
)
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
@@ -0,0 +1,182 @@
"""Manifest dataclasses for collect-changes.py output.
The manifest is the contract between the script and the review subagents.
Schema mirrors DESIGN.md.
"""
from __future__ import annotations
import json
from dataclasses import dataclass, field, asdict
from typing import Literal
Mode = Literal["local", "ref"]
Tool = Literal["terragrunt", "tofu"]
Action = Literal["create", "update", "delete", "replace", "read", "no-op", "unknown"]
Source = Literal["plan", "diff", "both"]
@dataclass
class InitResult:
ok: bool
stdout_tail: str
stderr_tail: str
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class PlanResult:
ok: bool
stdout_path: str
exit_code: int
summary: str
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class PlanUnit:
plan_dir: str
tool: Tool
init: InitResult
plan: PlanResult
triggered_by: list[str]
terragrunt_changed: bool = False
changed_files: list[str] = field(default_factory=list)
def to_dict(self) -> dict:
return {
"plan_dir": self.plan_dir,
"tool": self.tool,
"init": self.init.to_dict(),
"plan": self.plan.to_dict(),
"triggered_by": list(self.triggered_by),
"terragrunt_changed": self.terragrunt_changed,
"changed_files": list(self.changed_files),
}
@dataclass
class TrivyFinding:
check_id: str
title: str
severity: str
message: str
file: str
start_line: int = 0
end_line: int = 0
resource_type: str = ""
source: str = "trivy"
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class TflintFinding:
rule: str
severity: str
message: str
file: str
start_line: int = 0
end_line: int = 0
link: str = ""
source: str = "tflint"
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class CatalogInstance:
plan_dir: str
address_at_plan: str
action: Action
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class CatalogEntry:
source_dir: str
local_address: str
type: str
source: Source
instances: list[CatalogInstance] = field(default_factory=list)
block_header: str = ""
evidence_line: str = ""
key_attributes: dict[str, str | bool | int | list[str]] = field(default_factory=dict)
review_context: dict[str, object] = field(default_factory=dict)
block_file: str = ""
block_start: int = 0
block_end: int = 0
def to_dict(self) -> dict:
return {
"source_dir": self.source_dir,
"local_address": self.local_address,
"type": self.type,
"source": self.source,
"instances": [i.to_dict() for i in self.instances],
"block_header": self.block_header,
"evidence_line": self.evidence_line,
"key_attributes": dict(self.key_attributes),
"review_context": dict(self.review_context),
"block_file": self.block_file,
"block_start": self.block_start,
"block_end": self.block_end,
}
@dataclass
class ModuleGraphEntry:
callsites: list[str]
sibling_modules_at_callsites: list[str] = field(default_factory=list)
callsite_local_names: dict[str, dict[str, str]] = field(default_factory=dict)
def to_dict(self) -> dict:
return {
"callsites": list(self.callsites),
"sibling_modules_at_callsites": list(self.sibling_modules_at_callsites),
"callsite_local_names": {
callsite: dict(local_names)
for callsite, local_names in self.callsite_local_names.items()
},
}
@dataclass
class Manifest:
base_ref: str
head_ref: str
mode: Mode
default_branch: str
changed_source_dirs: list[str]
plan_units: list[PlanUnit]
catalog: list[CatalogEntry]
trivy_findings: list[TrivyFinding]
module_graph: dict[str, ModuleGraphEntry]
errors: list[str]
tflint_findings: list[TflintFinding] = field(default_factory=list)
def to_dict(self) -> dict:
return {
"base_ref": self.base_ref,
"head_ref": self.head_ref,
"mode": self.mode,
"default_branch": self.default_branch,
"changed_source_dirs": list(self.changed_source_dirs),
"plan_units": [p.to_dict() for p in self.plan_units],
"catalog": [c.to_dict() for c in self.catalog],
"trivy_findings": [f.to_dict() for f in self.trivy_findings],
"tflint_findings": [f.to_dict() for f in self.tflint_findings],
"module_graph": {k: v.to_dict() for k, v in self.module_graph.items()},
"errors": list(self.errors),
}
def to_json(self, indent: int = 2) -> str:
return json.dumps(self.to_dict(), indent=indent, sort_keys=False)
@@ -0,0 +1,90 @@
"""Repo-wide module callsite graph.
For each local-source module dir, list callsite dirs and sibling modules
(other modules instantiated alongside it at any callsite).
"""
from __future__ import annotations
from pathlib import Path, PurePosixPath
import hcl2
from scripts.manifest import ModuleGraphEntry
def _strip_hcl_quotes(s):
if isinstance(s, str) and len(s) >= 2 and s[0] == '"' and s[-1] == '"':
return s[1:-1]
return s
def _iter_module_blocks(tf_file: Path):
try:
with tf_file.open() as fh:
parsed = hcl2.load(fh)
except Exception:
return
for block in parsed.get("module", []):
if not isinstance(block, dict):
continue
for name, body in block.items():
if isinstance(body, dict):
src = body.get("source")
if isinstance(src, list):
src = src[0] if src else None
src = _strip_hcl_quotes(src)
yield _strip_hcl_quotes(name), src
def _is_local_source(src: str | None) -> bool:
if not isinstance(src, str):
return False
return src.startswith(("./", "../"))
def _resolve_local(callsite_dir: PurePosixPath, source: str) -> str:
combined = (callsite_dir / source).as_posix()
parts: list[str] = []
for part in combined.split("/"):
if part in ("", "."):
continue
if part == "..":
if parts:
parts.pop()
continue
parts.append(part)
return "/".join(parts)
def build_module_graph(repo_root: Path | str) -> dict[str, ModuleGraphEntry]:
root = Path(repo_root)
raw: dict[str, list[str]] = {}
callsite_to_modules: dict[str, dict[str, str]] = {}
for tf in root.rglob("*.tf"):
rel_dir = tf.parent.relative_to(root).as_posix()
callsite_dir = PurePosixPath(rel_dir)
for name, src in _iter_module_blocks(tf):
if not _is_local_source(src):
continue
target = _resolve_local(callsite_dir, src) # type: ignore[arg-type]
raw.setdefault(target, []).append(rel_dir)
callsite_to_modules.setdefault(rel_dir, {})[name] = target
out: dict[str, ModuleGraphEntry] = {}
for mod_dir, callsites in raw.items():
unique_callsites = sorted(set(callsites))
siblings: set[str] = set()
for cs in unique_callsites:
for other in callsite_to_modules.get(cs, {}).values():
if other != mod_dir:
siblings.add(other)
out[mod_dir] = ModuleGraphEntry(
callsites=unique_callsites,
sibling_modules_at_callsites=sorted(siblings),
callsite_local_names={
cs: dict(callsite_to_modules.get(cs, {}))
for cs in unique_callsites
},
)
return out
@@ -0,0 +1,59 @@
"""Parse terraform/tofu/terragrunt plan output for resource actions."""
from __future__ import annotations
import re
from dataclasses import dataclass
@dataclass(frozen=True)
class ParsedResource:
address: str
type: str
action: str
_HEADER_RE = re.compile(
r"^\s*#\s+(?P<addr>\S+)\s+(?:"
r"(?P<create>will be created)|"
r"(?P<update>will be updated in-place)|"
r"(?P<delete>will be destroyed)|"
r"(?P<replace>must be replaced|will be replaced)|"
r"(?P<read>will be read during apply)"
r")\s*$"
)
def _action_from_match(m: re.Match) -> str:
for k in ("create", "update", "delete", "replace", "read"):
if m.group(k):
return k
return "unknown"
def _type_from_address(address: str) -> str:
"""Strip module prefixes and any index suffix; return the resource type."""
parts = address.split(".")
i = 0
while i + 1 < len(parts) and parts[i].startswith("module"):
i += 2
if i < len(parts) and parts[i] == "data":
i += 1
if i >= len(parts):
return ""
type_token = parts[i]
return re.sub(r"\[.*\]$", "", type_token)
def parse_plan_resources(plan_stdout: str) -> list[ParsedResource]:
out: list[ParsedResource] = []
for line in plan_stdout.splitlines():
m = _HEADER_RE.match(line)
if not m:
continue
addr = m.group("addr")
out.append(ParsedResource(
address=addr,
type=_type_from_address(addr),
action=_action_from_match(m),
))
return out
@@ -0,0 +1,64 @@
"""Detect tool, run init + plan, capture output."""
from __future__ import annotations
import re
import subprocess
from pathlib import Path
from typing import Literal
from scripts.manifest import InitResult, PlanResult
Tool = Literal["terragrunt", "tofu"]
_TAIL_BYTES = 4096
_SUMMARY_RE = re.compile(
r"Plan:\s+(\d+\s+to add,\s*\d+\s+to change,\s*\d+\s+to destroy)\."
)
def detect_tool(plan_dir: Path | str) -> Tool:
p = Path(plan_dir)
if (p / "terragrunt.hcl").is_file():
return "terragrunt"
return "tofu"
def _tail(s: str, n: int = _TAIL_BYTES) -> str:
return s[-n:] if len(s) > n else s
def extract_summary(plan_stdout: str) -> str:
m = _SUMMARY_RE.search(plan_stdout)
if m:
return re.sub(r"\s+", " ", m.group(1)).strip()
if "No changes." in plan_stdout or "no changes." in plan_stdout.lower():
return "no changes"
return "unknown"
def run_init(plan_dir: Path | str, tool: Tool) -> InitResult:
args = [tool, "init", "-input=false", "-no-color"]
result = subprocess.run(
args, cwd=str(plan_dir), capture_output=True, text=True, check=False,
)
return InitResult(
ok=(result.returncode == 0),
stdout_tail=_tail(result.stdout),
stderr_tail=_tail(result.stderr),
)
def run_plan(plan_dir: Path | str, tool: Tool, stdout_path: Path) -> PlanResult:
args = [tool, "plan", "-input=false", "-no-color"]
result = subprocess.run(
args, cwd=str(plan_dir), capture_output=True, text=True, check=False,
)
stdout_path.parent.mkdir(parents=True, exist_ok=True)
stdout_path.write_text(result.stdout)
return PlanResult(
ok=(result.returncode == 0),
stdout_path=str(stdout_path),
exit_code=result.returncode,
summary=extract_summary(result.stdout),
)
@@ -0,0 +1,44 @@
"""Classify a directory as a plan unit, a reusable module, or unknown."""
from __future__ import annotations
from enum import Enum
from pathlib import Path
import hcl2
class DirKind(str, Enum):
PLAN_UNIT = "plan_unit"
MODULE = "module"
UNKNOWN = "unknown"
def _has_backend_block(tf_path: Path) -> bool:
try:
with tf_path.open() as fh:
parsed = hcl2.load(fh)
except Exception:
return False
for block in parsed.get("terraform", []):
if isinstance(block, dict) and "backend" in block:
return True
return False
def classify_dir(repo_root: Path | str, rel_dir: str) -> DirKind:
root = Path(repo_root)
full = root / rel_dir
if not full.is_dir():
raise FileNotFoundError(full)
if (full / "terragrunt.hcl").is_file():
return DirKind.PLAN_UNIT
tf_files = sorted(full.glob("*.tf")) + sorted(full.glob("*.tf.json"))
if not tf_files:
return DirKind.UNKNOWN
for tf in tf_files:
if tf.suffix == ".tf" and _has_backend_block(tf):
return DirKind.PLAN_UNIT
return DirKind.MODULE
@@ -0,0 +1,146 @@
"""Compute reference sets for the consistency reviewer."""
from __future__ import annotations
import re
from collections import Counter, defaultdict
from pathlib import Path
from scripts.manifest import ModuleGraphEntry
def _peer_dirs_at_depth(repo_root: Path, dir_path: str) -> list[str]:
parts = dir_path.split("/")
depth = len(parts)
parent = repo_root.joinpath(*parts[:-1]) if depth > 1 else repo_root
if not parent.is_dir():
return []
out: list[str] = []
for p in sorted(parent.iterdir()):
if not p.is_dir():
continue
rel = p.relative_to(repo_root).as_posix()
if rel != dir_path:
out.append(rel)
return out
def _terragrunt_refs(repo_root: Path, dir_path: str) -> list[str]:
"""Region peers + same-component cross-env (layout: live/<env>/<region>/<component>)."""
parts = dir_path.split("/")
if len(parts) != 4:
return _peer_dirs_at_depth(repo_root, dir_path)
live, env, region, component = parts
refs: set[str] = set()
region_dir = repo_root / live / env / region
if region_dir.is_dir():
for p in region_dir.iterdir():
if p.is_dir() and (p / "terragrunt.hcl").is_file():
rel = p.relative_to(repo_root).as_posix()
if rel != dir_path:
refs.add(rel)
live_dir = repo_root / live
if live_dir.is_dir():
for env_dir in live_dir.iterdir():
candidate = env_dir / region / component
if candidate.is_dir() and (candidate / "terragrunt.hcl").is_file():
rel = candidate.relative_to(repo_root).as_posix()
if rel != dir_path:
refs.add(rel)
env_live_dir = repo_root / live / env
if env_live_dir.is_dir():
for region_dir2 in env_live_dir.iterdir():
candidate = region_dir2 / component
if candidate.is_dir() and (candidate / "terragrunt.hcl").is_file():
rel = candidate.relative_to(repo_root).as_posix()
if rel != dir_path:
refs.add(rel)
return sorted(refs)
def compute_reference_sets(
repo_root: Path | str,
changed_dirs: set[str],
module_graph: dict[str, ModuleGraphEntry],
) -> dict[str, list[str]]:
root = Path(repo_root)
out: dict[str, list[str]] = {}
for d in sorted(changed_dirs):
if d.startswith("modules/"):
entry = module_graph.get(d)
out[d] = list(entry.sibling_modules_at_callsites) if entry else []
continue
full = root / d
if (full / "terragrunt.hcl").is_file():
out[d] = _terragrunt_refs(root, d)
continue
out[d] = _peer_dirs_at_depth(root, d)
return out
_ATTR_RE = re.compile(r"^\s*(?P<key>[A-Za-z0-9_]+)\s*=")
_MODULE_RE = re.compile(r'^\s*module\s+"(?P<name>[^"]+)"\s*\{')
def _iter_tf_lines(dir_path: Path) -> list[str]:
lines: list[str] = []
for tf in sorted(dir_path.glob("*.tf")):
lines.extend(tf.read_text().splitlines())
return lines
def _dir_signature(dir_path: Path) -> tuple[set[str], set[str]]:
attrs: set[str] = set()
modules: set[str] = set()
for line in _iter_tf_lines(dir_path):
attr_match = _ATTR_RE.match(line)
if attr_match:
attrs.add(attr_match.group("key"))
module_match = _MODULE_RE.match(line)
if module_match:
modules.add(module_match.group("name"))
return attrs, modules
def compute_consistency_norms(
repo_root: Path | str,
reference_sets: dict[str, list[str]],
) -> dict[str, dict[str, list[dict[str, object]]]]:
root = Path(repo_root)
out: dict[str, dict[str, list[dict[str, object]]]] = {}
for changed_dir, refs in sorted(reference_sets.items()):
attr_support: dict[str, list[str]] = defaultdict(list)
module_support: dict[str, list[str]] = defaultdict(list)
naming_tokens: Counter[str] = Counter()
usable_refs: list[str] = []
for ref in refs:
ref_path = root / ref
if not ref_path.is_dir():
continue
attrs, modules = _dir_signature(ref_path)
if not attrs and not modules:
continue
usable_refs.append(ref)
for attr in attrs:
attr_support[attr].append(ref)
for module in modules:
module_support[module].append(ref)
naming_tokens.update(Path(ref).parts[-1:])
out[changed_dir] = {
"attribute_norms": [
{"attribute": attr, "peer_dirs": sorted(peer_dirs)}
for attr, peer_dirs in sorted(attr_support.items())
if len(peer_dirs) >= 2
],
"module_wrapper_norms": [
{"module": module, "peer_dirs": sorted(peer_dirs)}
for module, peer_dirs in sorted(module_support.items())
if len(peer_dirs) >= 2
],
"naming_norms": [
{"token": token, "peer_dirs": sorted(usable_refs)}
for token, count in sorted(naming_tokens.items())
if count >= 2
],
}
return out
@@ -0,0 +1,95 @@
"""refresh-controls.py — populate data/controls/{fsbp,cis,meta}.json."""
from __future__ import annotations
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
import requests
_HERE = Path(__file__).resolve().parent
if str(_HERE.parent) not in sys.path:
sys.path.insert(0, str(_HERE.parent))
from scripts.controls_schema import Control, ControlsFile
from scripts.scrape_cis import CIS_URL, parse_cis_page
from scripts.scrape_fsbp import FSBP_INDEX_URL, parse_fsbp_index
_TIMEOUT = 30
_USER_AGENT = "audit-terraform-controls-refresh/1.0"
def _fetch(url: str) -> str:
r = requests.get(url, headers={"User-Agent": _USER_AGENT}, timeout=_TIMEOUT)
r.raise_for_status()
return r.text
def _now() -> datetime:
return datetime.now(timezone.utc)
def _build_fsbp() -> ControlsFile:
html = _fetch(FSBP_INDEX_URL)
rows = parse_fsbp_index(html, base_url=FSBP_INDEX_URL)
controls = [
Control(
control_id=f"FSBP {r.control_id}",
title=r.title,
severity=r.severity,
resource_types=r.resource_types,
requirement=r.requirement or r.title,
source_url=r.detail_url,
)
for r in rows
]
return ControlsFile(source="fsbp", fetched_at=_now(), controls=controls)
def _build_cis() -> ControlsFile:
html = _fetch(CIS_URL)
rows = parse_cis_page(html, source_url=CIS_URL)
controls = [
Control(
control_id=r.control_id,
title=r.title,
severity=r.severity,
resource_types=r.resource_types,
requirement=r.requirement,
source_url=r.source_url,
)
for r in rows
]
return ControlsFile(source="cis", fetched_at=_now(), controls=controls)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--output-dir", default=str(_HERE.parent / "data" / "controls"))
args = parser.parse_args(argv)
out_dir = Path(args.output_dir).resolve()
out_dir.mkdir(parents=True, exist_ok=True)
fsbp = _build_fsbp()
cis = _build_cis()
(out_dir / "fsbp.json").write_text(fsbp.to_json())
(out_dir / "cis.json").write_text(cis.to_json())
meta = {
"fsbp": {"url": FSBP_INDEX_URL, "fetched_at": fsbp.fetched_at.isoformat()},
"cis": {"url": CIS_URL, "fetched_at": cis.fetched_at.isoformat()},
}
(out_dir / "meta.json").write_text(json.dumps(meta, indent=2))
print(f"Wrote {len(fsbp.controls)} FSBP controls, {len(cis.controls)} CIS controls "
f"to {out_dir}")
return 0
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,41 @@
"""Map changed dirs to plan units, preserving the direct trigger directories."""
from __future__ import annotations
from pathlib import Path
from scripts.manifest import ModuleGraphEntry
from scripts.plan_unit import classify_dir, DirKind
def resolve_plan_units(
repo_root: Path | str,
changed_dirs: set[str],
module_graph: dict[str, ModuleGraphEntry],
) -> tuple[dict[str, list[str]], list[str]]:
"""Return (plan_units, orphan_modules).
plan_units maps plan_dir -> sorted-unique list of changed dirs that triggered it.
Direct Terragrunt or root-module changes remain in the trigger set so later
manifest assembly can attach per-plan-unit change context.
orphan_modules is a list of changed module dirs with zero callsites.
"""
plan_units: dict[str, set[str]] = {}
orphans: list[str] = []
for d in sorted(changed_dirs):
kind = classify_dir(repo_root, d)
if kind is DirKind.PLAN_UNIT:
plan_units.setdefault(d, set()).add(d)
elif kind is DirKind.MODULE:
entry = module_graph.get(d)
callsites = list(entry.callsites) if entry else []
if not callsites:
orphans.append(d)
continue
for cs in callsites:
plan_units.setdefault(cs, set()).add(d)
return (
{k: sorted(v) for k, v in sorted(plan_units.items())},
sorted(orphans),
)
@@ -0,0 +1,70 @@
"""Aggregate runs.jsonl into precision + token-cost stats."""
from __future__ import annotations
import json
import sys
from pathlib import Path
def compute_stats(log_path: Path) -> dict:
if not log_path.exists():
return {"by_agent": {}, "by_rule": {}, "runs": 0}
by_agent: dict[str, dict] = {}
by_rule: dict[str, dict] = {}
runs: set[str] = set()
for line in log_path.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
rec = json.loads(line)
agent = rec.get("agent", "?")
if rec["kind"] == "subagent_run":
runs.add(rec["run_id"])
a = by_agent.setdefault(agent, _empty_agent())
a["tokens"] += rec.get("input_tokens", 0) + rec.get("output_tokens", 0)
a["duration_ms"] += rec.get("duration_ms", 0)
a["runs"] += 1
elif rec["kind"] == "verdict":
verdict = rec["verdict"]
a = by_agent.setdefault(agent, _empty_agent())
a["total"] += 1
a[verdict] = a.get(verdict, 0) + 1
rule_key = f"{agent}/{rec['rule_id']}"
r = by_rule.setdefault(rule_key, _empty_rule())
r["total"] += 1
r[verdict] = r.get(verdict, 0) + 1
for a in by_agent.values():
a["precision"] = a["kept"] / a["total"] if a["total"] else 0.0
a["tokens_per_kept"] = a["tokens"] / a["kept"] if a["kept"] else float("inf")
for r in by_rule.values():
r["precision"] = r["kept"] / r["total"] if r["total"] else 0.0
return {"by_agent": by_agent, "by_rule": by_rule, "runs": len(runs)}
def _empty_agent() -> dict:
return {
"tokens": 0, "duration_ms": 0, "runs": 0,
"total": 0, "kept": 0, "dismissed": 0, "false_positive": 0,
}
def _empty_rule() -> dict:
return {"total": 0, "kept": 0, "dismissed": 0, "false_positive": 0}
def main(argv: list[str]) -> int:
log = Path.home() / ".claude/cache/audit-terraform/runs.jsonl"
stats = compute_stats(log)
print(json.dumps(stats, indent=2))
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
@@ -0,0 +1,117 @@
from __future__ import annotations
import json
import shutil
import subprocess
from pathlib import Path
from typing import Any
from scripts.manifest import TflintFinding, TrivyFinding
def _normalize_trivy_findings(payload: dict[str, Any]) -> list[TrivyFinding]:
findings: list[TrivyFinding] = []
for result in payload.get("Results", []):
for misconf in result.get("Misconfigurations", []):
cause = misconf.get("CauseMetadata") or {}
resource = cause.get("Resource") or ""
resource_type = ""
if isinstance(resource, str) and "." in resource:
resource_type = resource.split(".", 1)[0]
findings.append(TrivyFinding(
check_id=(
misconf.get("AVDID")
or misconf.get("ID")
or misconf.get("Query")
or ""
),
title=misconf.get("Title") or misconf.get("ID") or "",
severity=str(misconf.get("Severity") or "UNKNOWN").lower(),
message=misconf.get("Message") or misconf.get("Description") or "",
file=result.get("Target") or "",
start_line=int(cause.get("StartLine") or 0),
end_line=int(cause.get("EndLine") or 0),
resource_type=resource_type,
))
return findings
def _run_trivy_config(repo: Path, changed_files_in_repo: list[str]) -> tuple[dict[str, Any], list[TrivyFinding], str | None]:
changed_dirs = sorted({
str(Path(path).parent)
for path in changed_files_in_repo
})
scan_paths = [str((repo / rel).resolve()) for rel in changed_dirs] or [str(repo)]
result = subprocess.run(
["trivy", "config", "--quiet", "--format", "json", *scan_paths],
capture_output=True,
text=True,
check=False,
)
if result.returncode != 0:
stderr = result.stderr.strip() or result.stdout.strip()
return {}, [], f"trivy config failed: {stderr}"
try:
payload = json.loads(result.stdout) if result.stdout.strip() else {}
except json.JSONDecodeError as exc:
return {}, [], f"trivy config output was not valid JSON: {exc}"
return payload, _normalize_trivy_findings(payload), None
def _normalize_tflint_findings(payload: dict, *, scanned_dir: str) -> list[TflintFinding]:
findings: list[TflintFinding] = []
for issue in payload.get("issues", []) or []:
rule = issue.get("rule", {}) or {}
rng = issue.get("range", {}) or {}
start = rng.get("start", {}) or {}
end = rng.get("end", {}) or {}
rel_file = rng.get("filename") or ""
if "/" in rel_file:
full = rel_file
elif rel_file:
full = (Path(scanned_dir) / rel_file).as_posix()
else:
full = ""
findings.append(TflintFinding(
rule=rule.get("name") or "",
severity=str(rule.get("severity") or "warning").lower(),
message=issue.get("message") or "",
file=full,
start_line=int(start.get("line") or 0),
end_line=int(end.get("line") or 0),
link=rule.get("link") or "",
))
return findings
def _run_tflint(repo: Path, changed_files_in_repo: list[str]) -> tuple[dict, list[TflintFinding], str | None]:
if shutil.which("tflint") is None:
return {}, [], None
changed_dirs = sorted({
str(Path(p).parent.as_posix()) for p in changed_files_in_repo
if p.endswith((".tf", ".tf.json"))
})
if not changed_dirs:
return {}, [], None
aggregated: list[TflintFinding] = []
raw_payloads: dict[str, dict] = {}
errs: list[str] = []
for rel in changed_dirs:
target = (repo / rel).resolve()
if not target.is_dir():
continue
proc = subprocess.run(
["tflint", "--format", "json", "--chdir", str(target)],
capture_output=True, text=True, check=False,
)
if proc.returncode not in (0, 2):
errs.append(f"tflint failed in {rel}: {proc.stderr.strip() or proc.stdout.strip()}")
continue
try:
payload = json.loads(proc.stdout) if proc.stdout.strip() else {}
except json.JSONDecodeError as exc:
errs.append(f"tflint output in {rel} was not valid JSON: {exc}")
continue
raw_payloads[rel] = payload
aggregated.extend(_normalize_tflint_findings(payload, scanned_dir=rel))
return {"by_dir": raw_payloads}, aggregated, ("; ".join(errs) if errs else None)
@@ -0,0 +1,84 @@
"""Scrape the CIS AWS Foundations Benchmark page.
The page is a mapping table:
| Control ID and title (FSBP-style) | CIS v5.0.0 | CIS v3.0.0 | CIS v1.4.0 | CIS v1.2.0 |
We emit one CISRow per unified control. The `requirement` field summarises
which CIS versions reference it.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from bs4 import BeautifulSoup
from scripts.control_resource_map import control_id_to_types
CIS_URL = (
"https://docs.aws.amazon.com/securityhub/latest/userguide/"
"cis-aws-foundations-benchmark.html"
)
@dataclass
class CISRow:
control_id: str
title: str
severity: str
requirement: str
source_url: str
resource_types: list[str] = field(default_factory=list)
_TITLE_RE = re.compile(r"^\s*\[(?P<id>[A-Za-z][A-Za-z0-9]*\.\d+)\]\s*(?P<title>.+)$")
_VERSION_RE = re.compile(r"CIS\s+v(?P<ver>[\d.]+)", re.IGNORECASE)
def _extract_versions(headers: list[str]) -> list[str]:
"""From header cells, pull out version strings like '5.0.0' for each
column. Non-version columns yield empty string."""
versions: list[str] = []
for h in headers:
m = _VERSION_RE.search(h)
versions.append(m.group("ver") if m else "")
return versions
def parse_cis_page(html: str, source_url: str) -> list[CISRow]:
soup = BeautifulSoup(html, "html.parser")
rows: list[CISRow] = []
for table in soup.find_all("table"):
header_cells = table.find("tr").find_all(["th", "td"]) if table.find("tr") else []
header_texts = [c.get_text(strip=True) for c in header_cells]
versions = _extract_versions(header_texts)
if not any(versions):
continue
for tr in table.find_all("tr")[1:]:
cells = tr.find_all("td")
if len(cells) < 2:
continue
title_text = cells[0].get_text(strip=True)
m = _TITLE_RE.match(title_text)
if not m:
continue
unified_id = m.group("id")
title = m.group("title").strip()
version_refs: list[str] = []
for ver, cell in zip(versions[1:], cells[1:]):
if not ver:
continue
num = cell.get_text(strip=True)
if num:
version_refs.append(f"v{ver} §{num}")
requirement = "CIS " + ", ".join(version_refs) if version_refs else ""
rows.append(CISRow(
control_id=f"CIS {unified_id}",
title=title,
severity="medium",
requirement=requirement,
source_url=source_url,
resource_types=control_id_to_types(f"CIS {unified_id}"),
))
return rows
@@ -0,0 +1,61 @@
"""Scrape AWS Security Hub FSBP controls index page into structured rows.
The index page (https://docs.aws.amazon.com/securityhub/latest/userguide/fsbp-standard.html)
lists controls as <p><a href="./<svc>-controls.html#<id>">[<Service>.<Number>] <Title></a></p>.
This module only parses the index — severity defaults to "medium" because the
index doesn't surface severity. Detail-page enrichment is a future task.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from urllib.parse import urljoin
from bs4 import BeautifulSoup
from scripts.control_resource_map import control_id_to_types
FSBP_INDEX_URL = (
"https://docs.aws.amazon.com/securityhub/latest/userguide/"
"fsbp-standard.html"
)
@dataclass
class FSBPRow:
control_id: str
title: str
severity: str
detail_url: str
resource_types: list[str] = field(default_factory=list)
requirement: str = ""
_ANCHOR_RE = re.compile(r"^\s*\[(?P<id>[A-Za-z][A-Za-z0-9]*\.\d+)\]\s*(?P<title>.+)$")
def parse_fsbp_index(html: str, base_url: str) -> list[FSBPRow]:
soup = BeautifulSoup(html, "html.parser")
rows: list[FSBPRow] = []
seen: set[str] = set()
for a in soup.find_all("a"):
text = a.get_text(strip=True)
m = _ANCHOR_RE.match(text)
if not m:
continue
control_id = m.group("id")
if control_id in seen:
continue
seen.add(control_id)
title = m.group("title").strip()
href = a.get("href", "")
detail_url = urljoin(base_url, href) if href else ""
rows.append(FSBPRow(
control_id=control_id,
title=title,
severity="medium",
detail_url=detail_url,
resource_types=control_id_to_types(control_id),
))
return rows
@@ -0,0 +1,64 @@
from __future__ import annotations
def _slice_for_agent(manifest_dict: dict, agent: str) -> dict:
"""Return a subset of the manifest appropriate to the named agent.
- fsbp/cis/aws-bp: AWS-filtered catalog + plan_units summary + trivy + tflint + errors
- consistency: full catalog + full plan_units + trivy + tflint + errors
- tf-hygiene: full catalog + full plan_units + tflint + changed_source_dirs + errors (no trivy)
- walkthrough: full catalog + plan_units summary + changed_source_dirs + errors (no findings)
"""
base = {
"base_ref": manifest_dict["base_ref"],
"head_ref": manifest_dict["head_ref"],
"mode": manifest_dict["mode"],
"default_branch": manifest_dict["default_branch"],
"errors": list(manifest_dict["errors"]),
}
catalog = manifest_dict["catalog"]
if agent in ("fsbp", "cis", "aws-bp"):
return {
**base,
"catalog": [e for e in catalog if e["type"].startswith("aws_")],
"trivy_findings": list(manifest_dict.get("trivy_findings", [])),
"tflint_findings": list(manifest_dict.get("tflint_findings", [])),
"plan_units": [
{"plan_dir": pu["plan_dir"], "tool": pu["tool"],
"plan_ok": pu["plan"]["ok"], "summary": pu["plan"]["summary"],
"terragrunt_changed": pu["terragrunt_changed"],
"changed_files": list(pu["changed_files"])}
for pu in manifest_dict["plan_units"]
],
}
if agent == "consistency":
return {
**base,
"changed_source_dirs": list(manifest_dict["changed_source_dirs"]),
"catalog": catalog,
"trivy_findings": list(manifest_dict.get("trivy_findings", [])),
"tflint_findings": list(manifest_dict.get("tflint_findings", [])),
"plan_units": [dict(pu) for pu in manifest_dict["plan_units"]],
}
if agent == "walkthrough":
return {
**base,
"changed_source_dirs": list(manifest_dict["changed_source_dirs"]),
"catalog": catalog,
"plan_units": [
{"plan_dir": pu["plan_dir"], "tool": pu["tool"],
"plan_ok": pu["plan"]["ok"], "summary": pu["plan"]["summary"],
"terragrunt_changed": pu["terragrunt_changed"],
"changed_files": list(pu["changed_files"])}
for pu in manifest_dict["plan_units"]
],
}
if agent == "tf-hygiene":
return {
**base,
"changed_source_dirs": list(manifest_dict["changed_source_dirs"]),
"catalog": catalog,
"tflint_findings": list(manifest_dict.get("tflint_findings", [])),
"plan_units": [dict(pu) for pu in manifest_dict["plan_units"]],
}
return manifest_dict
@@ -0,0 +1,220 @@
"""Locate a specific resource block in a directory's .tf files."""
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import re
from scripts.hcl_diff import find_resource_blocks
_KEY_ATTRIBUTES = {
"kms_key_id",
"kms_key_arn",
"bucket_key_enabled",
"publicly_accessible",
"acl",
"versioning",
"tags",
"deletion_protection",
}
_VAR_REF_RE = re.compile(r"\bvar\.([A-Za-z0-9_]+)")
_LOCAL_REF_RE = re.compile(r"\blocal\.([A-Za-z0-9_]+)")
_DATA_REF_RE = re.compile(r"\bdata\.aws_iam_policy_document\.([A-Za-z0-9_]+)")
_SG_REF_TEMPLATE = 'security_group_id = aws_security_group.{name}.id'
@dataclass(frozen=True)
class BlockLocation:
file: Path
start_line: int
end_line: int
text: str
header: str
evidence_line: str
key_attributes: dict[str, str | bool | int | list[str]]
review_context: dict[str, object]
def _parse_value(raw: str) -> str | bool | int | list[str]:
value = raw.strip().rstrip(",")
if value.lower() in {"true", "false"}:
return value.lower() == "true"
if re.fullmatch(r"-?\d+", value):
return int(value)
if value.startswith('"') and value.endswith('"'):
return value[1:-1]
if value.startswith("[") and value.endswith("]"):
inner = value[1:-1].strip()
if not inner:
return []
parts = [part.strip().strip('"') for part in inner.split(",")]
return [part for part in parts if part]
return value
def _extract_key_attributes(lines: list[str]) -> dict[str, str | bool | int | list[str]]:
out: dict[str, str | bool | int | list[str]] = {}
for line in lines:
stripped = line.strip()
if "=" not in stripped or stripped.startswith("#"):
continue
key, _, raw_value = stripped.partition("=")
key = key.strip()
if key not in _KEY_ATTRIBUTES:
continue
out[key] = _parse_value(raw_value)
return out
def _find_named_block(lines: list[str], header_re: re.Pattern[str], name: str) -> list[str]:
i = 0
while i < len(lines):
line = lines[i]
match = header_re.match(line)
if not match or match.group("name") != name:
i += 1
continue
depth = line.count("{") - line.count("}")
j = i
while depth > 0 and j + 1 < len(lines):
j += 1
depth += lines[j].count("{") - lines[j].count("}")
return lines[i:j + 1]
return []
def _collect_variable_defaults(root: Path, names: set[str]) -> dict[str, str | bool | int | list[str]]:
out: dict[str, str | bool | int | list[str]] = {}
header_re = re.compile(r'^\s*variable\s+"(?P<name>[^"]+)"\s*\{')
for tf in sorted(root.glob("*.tf")):
lines = tf.read_text().splitlines()
for name in names:
if name in out:
continue
block = _find_named_block(lines, header_re, name)
for line in block[1:]:
stripped = line.strip()
if stripped.startswith("default"):
_, _, raw = stripped.partition("=")
out[name] = _parse_value(raw)
break
return out
def _collect_locals(root: Path, names: set[str]) -> dict[str, str | bool | int | list[str]]:
out: dict[str, str | bool | int | list[str]] = {}
header_re = re.compile(r"^\s*locals\s*\{")
for tf in sorted(root.glob("*.tf")):
lines = tf.read_text().splitlines()
i = 0
while i < len(lines):
line = lines[i]
if not header_re.match(line):
i += 1
continue
depth = line.count("{") - line.count("}")
j = i
while depth > 0 and j + 1 < len(lines):
j += 1
depth += lines[j].count("{") - lines[j].count("}")
for block_line in lines[i + 1:j]:
stripped = block_line.strip()
if "=" not in stripped:
continue
key, _, raw = stripped.partition("=")
key = key.strip()
if key in names and key not in out:
out[key] = _parse_value(raw)
i = j + 1
return out
def _collect_related_policy_docs(root: Path, names: set[str]) -> list[str]:
out: list[str] = []
header_re = re.compile(
r'^\s*data\s+"aws_iam_policy_document"\s+"(?P<name>[^"]+)"\s*\{'
)
for tf in sorted(root.glob("*.tf")):
lines = tf.read_text().splitlines()
for name in names:
block = _find_named_block(lines, header_re, name)
if block:
out.append(block[0].strip())
return out
def _collect_related_sg_rules(root: Path, name: str) -> list[str]:
out: list[str] = []
target = _SG_REF_TEMPLATE.format(name=name)
header_re = re.compile(
r'^\s*resource\s+"aws_security_group_rule"\s+"(?P<name>[^"]+)"\s*\{'
)
for tf in sorted(root.glob("*.tf")):
lines = tf.read_text().splitlines()
i = 0
while i < len(lines):
line = lines[i]
if not header_re.match(line):
i += 1
continue
depth = line.count("{") - line.count("}")
j = i
while depth > 0 and j + 1 < len(lines):
j += 1
depth += lines[j].count("{") - lines[j].count("}")
block = lines[i:j + 1]
if any(target in block_line for block_line in block[1:]):
out.append(block[0].strip())
i = j + 1
return out
def _build_review_context(root: Path, rtype: str, rname: str, block_lines: list[str]) -> dict[str, object]:
joined = "\n".join(block_lines)
variable_names = set(_VAR_REF_RE.findall(joined))
local_names = set(_LOCAL_REF_RE.findall(joined))
policy_doc_names = set(_DATA_REF_RE.findall(joined))
related_blocks: list[str] = []
related_blocks.extend(_collect_related_policy_docs(root, policy_doc_names))
if rtype == "aws_security_group":
related_blocks.extend(_collect_related_sg_rules(root, rname))
return {
"variables": _collect_variable_defaults(root, variable_names),
"locals": _collect_locals(root, local_names),
"related_blocks": related_blocks,
}
def find_block(source_dir: Path | str, local_address: str) -> BlockLocation | None:
"""Search .tf files in `source_dir` for a resource matching `local_address`.
`local_address` is `<type>.<name>` (e.g. `aws_iam_role.svc`). Returns the
first match found. Returns None if no match.
"""
if "." not in local_address:
return None
rtype, _, rname = local_address.partition(".")
root = Path(source_dir)
for tf in sorted(root.glob("*.tf")):
for blk in find_resource_blocks(tf):
if blk.type == rtype and blk.name == rname:
lines = tf.read_text().splitlines()
block_lines = lines[blk.start - 1: blk.end]
header = block_lines[0] if block_lines else ""
evidence_line = header
for line in block_lines[1:]:
if line.strip():
evidence_line = line.strip()
break
return BlockLocation(
file=tf,
start_line=blk.start,
end_line=blk.end,
text="\n".join(block_lines),
header=header,
evidence_line=evidence_line,
key_attributes=_extract_key_attributes(block_lines[1:]),
review_context=_build_review_context(root, rtype, rname, block_lines),
)
return None
@@ -0,0 +1,53 @@
"""Append-only JSONL telemetry for audit-terraform runs."""
from __future__ import annotations
import json
from datetime import datetime, timezone
from pathlib import Path
def _now() -> str:
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def _append(log_path: Path, record: dict) -> None:
log_path.parent.mkdir(parents=True, exist_ok=True)
with log_path.open("a", encoding="utf-8") as f:
f.write(json.dumps(record) + "\n")
def append_subagent_run(
log_path: Path, *, run_id: str, repo: str, mode: str, agent: str,
model: str, input_tokens: int, output_tokens: int,
duration_ms: int, finding_count: int,
) -> None:
_append(log_path, {
"kind": "subagent_run", "ts": _now(),
"run_id": run_id, "repo": repo, "mode": mode, "agent": agent,
"model": model,
"input_tokens": input_tokens, "output_tokens": output_tokens,
"duration_ms": duration_ms, "finding_count": finding_count,
})
def append_verdict(
log_path: Path, *, run_id: str, agent: str, rule_id: str,
file: str, line: int, verdict: str, notes: str = "",
) -> None:
if verdict not in {"kept", "dismissed", "false_positive"}:
raise ValueError(f"invalid verdict: {verdict!r}")
_append(log_path, {
"kind": "verdict", "ts": _now(),
"run_id": run_id, "agent": agent, "rule_id": rule_id,
"file": file, "line": line, "verdict": verdict, "notes": notes,
})
def read_runs(log_path: Path) -> list[dict]:
if not log_path.exists():
return []
return [
json.loads(line)
for line in log_path.read_text(encoding="utf-8").splitlines()
if line.strip()
]
@@ -0,0 +1,5 @@
import sys
from pathlib import Path
_root = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(_root))
@@ -0,0 +1,41 @@
<!DOCTYPE html>
<html><body>
<h2>CIS AWS Foundations Benchmark version 5.0.0</h2>
<table>
<tr>
<th>Control ID and title</th>
<th>CIS v5.0.0 requirement</th>
<th>CIS v3.0.0 requirement</th>
<th>CIS v1.4.0 requirement</th>
<th>CIS v1.2.0 requirement</th>
</tr>
<tr>
<td>[Account.1] Security contact information should be provided for an AWS account</td>
<td>1.2</td>
<td>1.2</td>
<td>1.2</td>
<td>1.18</td>
</tr>
<tr>
<td>[CloudTrail.1] CloudTrail should be enabled and configured with at least one multi-Region trail</td>
<td>3.1</td>
<td>3.1</td>
<td>3.1</td>
<td>2.1</td>
</tr>
<tr>
<td>[IAM.5] MFA should be enabled for all IAM users that have a console password</td>
<td>1.10</td>
<td>1.10</td>
<td></td>
<td>1.2</td>
</tr>
<tr>
<td>[S3.5] S3 general purpose buckets should require requests to use SSL</td>
<td>2.1.2</td>
<td>2.1.2</td>
<td>2.1.2</td>
<td></td>
</tr>
</table>
</body></html>
@@ -0,0 +1,22 @@
<!DOCTYPE html>
<html><body>
<h1>AWS Foundational Security Best Practices standard</h1>
<h3>S3 controls</h3>
<p>
<a href="./s3-controls.html#s3-1">[S3.1] S3 general purpose buckets should have block public access settings enabled</a>
</p>
<p>
<a href="./s3-controls.html#s3-5">[S3.5] S3 general purpose buckets should require requests to use SSL</a>
</p>
<h3>IAM controls</h3>
<p>
<a href="./iam-controls.html#iam-5">[IAM.5] MFA should be enabled for all IAM users that have a console password</a>
</p>
<h3>Other</h3>
<p>
<a href="./account-controls.html#account-1">[Account.1] Security contact information should be provided for an AWS account</a>
</p>
</body></html>
@@ -0,0 +1,35 @@
OpenTofu used the selected providers to generate the following execution
plan. Resource actions are indicated with the following symbols:
+ create
~ update in-place
- destroy
-/+ destroy and then create replacement
OpenTofu will perform the following actions:
# aws_s3_bucket.audit_logs will be created
+ resource "aws_s3_bucket" "audit_logs" {
+ bucket = "audit-logs-prod"
}
# module.app_role.aws_iam_role.this will be updated in-place
~ resource "aws_iam_role" "this" {
name = "app-role"
~ assume_role_policy = jsonencode(
~ {
...
}
)
}
# aws_security_group.legacy will be destroyed
- resource "aws_security_group" "legacy" {
- id = "sg-123" -> null
}
# aws_iam_user.svc must be replaced
-/+ resource "aws_iam_user" "svc" {
~ name = "old" -> "new" # forces replacement
}
Plan: 2 to add, 1 to change, 1 to destroy.
@@ -0,0 +1,22 @@
terraform {
required_version = ">= 1.5"
backend "local" {}
required_providers {
null = { source = "hashicorp/null", version = "~> 3.2" }
random = { source = "hashicorp/random", version = "~> 3.6" }
}
}
module "widget_a" {
source = "./modules/widget"
name = "alpha"
}
module "widget_b" {
source = "./modules/widget"
name = "beta"
}
resource "null_resource" "top" {
triggers = { ts = "1" }
}
@@ -0,0 +1,8 @@
resource "null_resource" "thing" {
triggers = { name = var.name }
}
resource "random_string" "id" {
length = 8
special = false
}
@@ -0,0 +1,3 @@
variable "name" {
type = string
}
@@ -0,0 +1,24 @@
from pathlib import Path
SKILL_ROOT = Path(__file__).resolve().parent.parent
def test_tf_hygiene_agent_prompt_exists():
p = SKILL_ROOT / "agents" / "tf-hygiene-reviewer.md"
assert p.exists(), f"missing agent prompt at {p}"
text = p.read_text()
assert "tf-hygiene-reviewer" in text
assert "OUTPUT" in text
assert "MANIFEST" in text
assert "tflint_findings" in text
def test_tf_hygiene_agent_disclaims_security_lane():
text = (SKILL_ROOT / "agents" / "tf-hygiene-reviewer.md").read_text().lower()
assert "security" in text
assert "aws-bp" in text or "not security" in text
def test_tf_hygiene_agent_disclaims_consistency_lane():
text = (SKILL_ROOT / "agents" / "tf-hygiene-reviewer.md").read_text().lower()
assert "consistency" in text
@@ -0,0 +1,153 @@
from scripts.catalog import build_catalog, PlanHit, DiffHit
from scripts.manifest import ModuleGraphEntry
def test_plan_and_diff_match_yields_both():
plan_hits = {
"live/prod/app": [
PlanHit(address="aws_s3_bucket.x", type="aws_s3_bucket", action="create"),
],
}
diff_hits = [
DiffHit(source_dir="live/prod/app",
local_address="aws_s3_bucket.x", type="aws_s3_bucket"),
]
entries = build_catalog(plan_hits, diff_hits, module_graph={})
assert len(entries) == 1
e = entries[0]
assert e.source == "both"
assert e.source_dir == "live/prod/app"
assert e.local_address == "aws_s3_bucket.x"
assert len(e.instances) == 1
assert e.instances[0].plan_dir == "live/prod/app"
assert e.instances[0].address_at_plan == "aws_s3_bucket.x"
def test_plan_only_is_drift():
plan_hits = {
"live/prod/app": [
PlanHit(address="aws_s3_bucket.x", type="aws_s3_bucket", action="update"),
],
}
entries = build_catalog(plan_hits, diff_hits=[], module_graph={})
assert len(entries) == 1
assert entries[0].source == "plan"
assert entries[0].instances[0].action == "update"
def test_module_callsites_collapse_to_single_entry():
graph = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app", "live/staging/app"],
)
}
plan_hits = {
"live/prod/app": [
PlanHit(address="module.role.aws_iam_role.this",
type="aws_iam_role", action="update"),
],
"live/staging/app": [
PlanHit(address="module.role.aws_iam_role.this",
type="aws_iam_role", action="update"),
],
}
entries = build_catalog(plan_hits, diff_hits=[], module_graph=graph)
assert len(entries) == 1
e = entries[0]
assert e.source_dir == "modules/app-role"
assert e.local_address == "aws_iam_role.this"
assert {i.plan_dir for i in e.instances} == {"live/prod/app", "live/staging/app"}
def test_plan_module_addresses_resolve_to_matching_local_module():
graph = {
"modules/app-policy": ModuleGraphEntry(
callsites=["live/prod/app"],
callsite_local_names={
"live/prod/app": {
"policy": "modules/app-policy",
"role": "modules/app-role",
}
},
),
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app"],
callsite_local_names={
"live/prod/app": {
"policy": "modules/app-policy",
"role": "modules/app-role",
}
},
),
}
plan_hits = {
"live/prod/app": [
PlanHit(address="module.role.aws_iam_role.this",
type="aws_iam_role", action="update"),
PlanHit(address="module.policy.aws_iam_policy.this",
type="aws_iam_policy", action="create"),
],
}
entries = build_catalog(plan_hits, diff_hits=[], module_graph=graph)
by_key = {(e.source_dir, e.local_address): e for e in entries}
assert ("modules/app-role", "aws_iam_role.this") in by_key
assert ("modules/app-policy", "aws_iam_policy.this") in by_key
assert ("modules/app-role", "aws_iam_policy.this") not in by_key
assert ("modules/app-policy", "aws_iam_role.this") not in by_key
def test_plan_module_resolution_does_not_guess_when_local_name_missing():
graph = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app"],
callsite_local_names={
"live/prod/app": {
"role": "modules/app-role",
}
},
),
}
plan_hits = {
"live/prod/app": [
PlanHit(address="module.policy.aws_iam_policy.this",
type="aws_iam_policy", action="create"),
],
}
entries = build_catalog(plan_hits, diff_hits=[], module_graph=graph)
assert len(entries) == 1
assert entries[0].source_dir == "live/prod/app"
assert entries[0].local_address == "aws_iam_policy.this"
def test_diff_only_module_fans_out_instances_to_callsites():
graph = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app", "live/staging/app"],
)
}
diff_hits = [
DiffHit(source_dir="modules/app-role",
local_address="aws_iam_role.this", type="aws_iam_role"),
]
entries = build_catalog(plan_hits={}, diff_hits=diff_hits, module_graph=graph)
assert len(entries) == 1
e = entries[0]
assert e.source == "diff"
assert {i.plan_dir for i in e.instances} == {"live/prod/app", "live/staging/app"}
def test_diff_only_module_with_no_callsites_emits_one_instance():
diff_hits = [
DiffHit(source_dir="modules/lonely",
local_address="aws_iam_role.this", type="aws_iam_role"),
]
entries = build_catalog(plan_hits={}, diff_hits=diff_hits, module_graph={})
assert len(entries) == 1
e = entries[0]
assert e.source == "diff"
assert len(e.instances) == 1
assert e.instances[0].plan_dir == "modules/lonely"
@@ -0,0 +1,421 @@
import importlib.util
import json
import shutil
import subprocess
from pathlib import Path
from unittest.mock import patch, MagicMock
import pytest
_SKILL_ROOT = Path(__file__).resolve().parent.parent
def _load_cli():
spec = importlib.util.spec_from_file_location(
"collect_changes", _SKILL_ROOT / "scripts" / "collect-changes.py"
)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def _stage_fixture(dest: Path) -> Path:
src = _SKILL_ROOT / "tests" / "fixtures" / "tofu-sample"
shutil.copytree(src, dest)
return dest
def _fake_subprocess(cmd, **kwargs):
if cmd[:2] == ["git", "-C"]:
sub = cmd[2:]
else:
sub = cmd
if "symbolic-ref" in sub:
return MagicMock(returncode=0, stdout="refs/remotes/origin/main\n", stderr="")
if "fetch" in sub:
return MagicMock(returncode=0, stdout="", stderr="")
if "rev-parse" in sub and "--verify" in sub:
return MagicMock(returncode=0, stdout="abc123\n", stderr="")
if cmd[:1] == ["git"] and "diff" in cmd:
diff = (
"diff --git a/main.tf b/main.tf\n"
"index 1..2 100644\n"
"--- a/main.tf\n"
"+++ b/main.tf\n"
"@@ -20,0 +21,1 @@\n"
"+ bucket_prefix = \"x\"\n"
"diff --git a/modules/widget/main.tf b/modules/widget/main.tf\n"
"index 3..4 100644\n"
"--- a/modules/widget/main.tf\n"
"+++ b/modules/widget/main.tf\n"
"@@ -1,0 +2,1 @@\n"
"+ # touched\n"
)
return MagicMock(returncode=0, stdout=diff, stderr="")
if cmd[0] == "tofu" and "init" in cmd:
return MagicMock(returncode=0, stdout="initialized\n", stderr="")
if cmd[0] == "tofu" and "plan" in cmd:
plan = (
"OpenTofu will perform the following actions:\n\n"
" # null_resource.top will be updated in-place\n"
" ~ resource \"null_resource\" \"top\" {}\n\n"
" # module.widget_a.null_resource.thing will be updated in-place\n"
" ~ resource \"null_resource\" \"thing\" {}\n\n"
" # module.widget_b.null_resource.thing will be updated in-place\n"
" ~ resource \"null_resource\" \"thing\" {}\n\n"
"Plan: 0 to add, 3 to change, 0 to destroy.\n"
)
return MagicMock(returncode=0, stdout=plan, stderr="")
if cmd[:3] == ["trivy", "config", "--quiet"]:
trivy = {
"SchemaVersion": 2,
"Results": [
{
"Target": "main.tf",
"Class": "config",
"Type": "terraform",
"Misconfigurations": [
{
"ID": "AVD-AWS-0089",
"AVDID": "AVD-AWS-0089",
"Title": "S3 bucket allows public ACL",
"Description": "Buckets should not allow public ACLs.",
"Message": "Bucket ACL allows public access.",
"Severity": "HIGH",
"CauseMetadata": {
"Resource": "aws_s3_bucket.audit_logs",
"StartLine": 21,
"EndLine": 30,
},
}
],
}
],
}
return MagicMock(returncode=0, stdout=json.dumps(trivy), stderr="")
return MagicMock(returncode=0, stdout="", stderr="")
def test_cli_happy_path(tmp_path):
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
with patch("subprocess.run", side_effect=_fake_subprocess):
rc = mod.main([
"--repo", str(repo),
"--base", "main",
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0, (out_dir / "manifest.json").read_text()
manifest = json.loads((out_dir / "manifest.json").read_text())
assert manifest["mode"] == "local"
assert manifest["base_ref"] == "main"
assert set(manifest["changed_source_dirs"]) == {".", "modules/widget"}
plan_dirs = {pu["plan_dir"] for pu in manifest["plan_units"]}
assert plan_dirs == {"."}
pu = manifest["plan_units"][0]
assert pu["tool"] == "tofu"
assert pu["plan"]["summary"] == "0 to add, 3 to change, 0 to destroy"
catalog = manifest["catalog"]
keys = {(e["source_dir"], e["local_address"]) for e in catalog}
assert (".", "null_resource.top") in keys
assert ("modules/widget", "null_resource.thing") in keys
module_entries = [e for e in catalog if e["source_dir"] == "modules/widget"]
for e in module_entries:
plan_dirs = {i["plan_dir"] for i in e["instances"]}
assert plan_dirs == {"."}, e
addrs = {i["address_at_plan"] for i in e["instances"]}
assert any(a.startswith("module.widget_a") for a in addrs)
assert any(a.startswith("module.widget_b") for a in addrs)
for e in module_entries:
assert e["block_header"], f"missing block_header on {e}"
assert e["evidence_line"], f"missing evidence_line on {e}"
assert isinstance(e["key_attributes"], dict)
assert "review_context" in e
assert set(e["review_context"]) == {"variables", "locals", "related_blocks"}
assert e["block_file"].endswith(".tf")
assert e["block_start"] >= 1
assert "block_text" not in e
assert (out_dir / "trivy-findings.json").is_file()
assert manifest["trivy_findings"] == [
{
"check_id": "AVD-AWS-0089",
"title": "S3 bucket allows public ACL",
"severity": "high",
"message": "Bucket ACL allows public access.",
"file": "main.tf",
"start_line": 21,
"end_line": 30,
"resource_type": "aws_s3_bucket",
"source": "trivy",
}
]
refs = json.loads((out_dir / "reference_sets.json").read_text())
assert "." in refs or "modules/widget" in refs
def test_cli_writes_per_agent_slices(tmp_path):
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
with patch("subprocess.run", side_effect=_fake_subprocess):
rc = mod.main([
"--repo", str(repo), "--base", "main", "--head", "HEAD",
"--output-dir", str(out_dir), "--mode", "local",
])
assert rc == 0
full = json.loads((out_dir / "manifest.json").read_text())
for agent in ("fsbp", "cis", "aws-bp", "consistency"):
sliced = json.loads((out_dir / f"manifest-{agent}.json").read_text())
assert sliced["base_ref"] == full["base_ref"]
if agent in ("fsbp", "cis", "aws-bp"):
for entry in sliced["catalog"]:
assert entry["type"].startswith("aws_"), (
f"non-aws type leaked into {agent}: {entry['type']}"
)
assert sliced["catalog"] == []
assert sliced["trivy_findings"] == full["trivy_findings"]
if agent == "consistency":
assert sliced["changed_source_dirs"] == full["changed_source_dirs"]
assert len(sliced["catalog"]) == len(full["catalog"])
assert sliced["trivy_findings"] == full["trivy_findings"]
def test_cli_preserves_terragrunt_only_change_context(tmp_path):
repo = _stage_fixture(tmp_path / "repo")
terragrunt_dir = repo / "live" / "prod" / "app"
terragrunt_dir.mkdir(parents=True)
(terragrunt_dir / "terragrunt.hcl").write_text("inputs = { instance_count = 2 }\n")
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
diff = (
"diff --git a/live/prod/app/terragrunt.hcl b/live/prod/app/terragrunt.hcl\n"
"index 1..2 100644\n"
"--- a/live/prod/app/terragrunt.hcl\n"
"+++ b/live/prod/app/terragrunt.hcl\n"
"@@ -1 +1 @@\n"
"-inputs = { instance_count = 1 }\n"
"+inputs = { instance_count = 2 }\n"
)
return MagicMock(returncode=0, stdout=diff, stderr="")
if cmd[0] == "terragrunt" and "init" in cmd:
return MagicMock(returncode=0, stdout="initialized\n", stderr="")
if cmd[0] == "terragrunt" and "plan" in cmd:
plan = (
"OpenTofu will perform the following actions:\n\n"
"Plan: 0 to add, 0 to change, 0 to destroy.\n"
)
return MagicMock(returncode=0, stdout=plan, stderr="")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
rc = mod.main([
"--repo", str(repo),
"--base", "main",
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0, (out_dir / "manifest.json").read_text()
manifest = json.loads((out_dir / "manifest.json").read_text())
assert "live/prod/app" in manifest["changed_source_dirs"]
assert manifest["catalog"] == []
assert manifest["plan_units"] == [{
"plan_dir": "live/prod/app",
"tool": "terragrunt",
"init": {"ok": True, "stdout_tail": "initialized\n", "stderr_tail": ""},
"plan": {
"ok": True,
"stdout_path": str(out_dir / "plans" / "live_prod_app.txt"),
"exit_code": 0,
"summary": "0 to add, 0 to change, 0 to destroy",
},
"triggered_by": ["live/prod/app"],
"terragrunt_changed": True,
"changed_files": ["live/prod/app/terragrunt.hcl"],
}]
fsbp = json.loads((out_dir / "manifest-fsbp.json").read_text())
assert fsbp["plan_units"] == [{
"plan_dir": "live/prod/app",
"tool": "terragrunt",
"plan_ok": True,
"summary": "0 to add, 0 to change, 0 to destroy",
"terragrunt_changed": True,
"changed_files": ["live/prod/app/terragrunt.hcl"],
}]
consistency = json.loads((out_dir / "manifest-consistency.json").read_text())
assert consistency["plan_units"] == manifest["plan_units"]
def test_cli_aborts_when_plan_fails(tmp_path):
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[0] == "tofu" and "plan" in cmd:
return MagicMock(
returncode=1,
stdout="Error: syntax error in main.tf\n",
stderr="",
)
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
rc = mod.main([
"--repo", str(repo),
"--base", "main",
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 1
manifest = json.loads((out_dir / "manifest.json").read_text())
assert any("syntax error" in e for e in manifest["errors"])
def test_cli_records_git_diff_failure_in_manifest(tmp_path):
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
return MagicMock(returncode=128, stdout="",
stderr="fatal: bad revision 'main...HEAD'\n")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
rc = mod.main([
"--repo", str(repo),
"--base", "main",
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 1
manifest = json.loads((out_dir / "manifest.json").read_text())
assert any("bad revision" in e for e in manifest["errors"])
def test_cli_auto_resolves_base_to_origin(tmp_path):
"""With no --base, the CLI fetches and diffs against origin/<default>."""
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
captured_diff_base: list[str] = []
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
for arg in cmd:
if "..." in arg:
captured_diff_base.append(arg)
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0
assert captured_diff_base, "git diff was never invoked"
assert captured_diff_base[0].startswith("origin/main..."), (
f"diff base should be origin/<default>, got: {captured_diff_base[0]}"
)
manifest = json.loads((out_dir / "manifest.json").read_text())
assert manifest["base_ref"] == "origin/main"
def test_cli_falls_back_to_local_branch_when_origin_missing(tmp_path):
"""If origin/<default> doesn't exist (no remote, fetch fails), fall back."""
repo = _stage_fixture(tmp_path / "repo")
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:2] == ["git", "-C"]:
sub = cmd[2:]
else:
sub = cmd
if "rev-parse" in sub and "--verify" in sub:
return MagicMock(returncode=1, stdout="", stderr="fatal: unknown revision\n")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0
manifest = json.loads((out_dir / "manifest.json").read_text())
assert manifest["base_ref"] == "main"
def _has_tofu() -> bool:
from shutil import which
return which("tofu") is not None
@pytest.mark.integration
@pytest.mark.skipif(not _has_tofu(), reason="tofu not installed")
def test_cli_real_tofu_plan(tmp_path, monkeypatch):
repo = _stage_fixture(tmp_path / "repo")
subprocess.run(["git", "init", "-q", str(repo)], check=True)
subprocess.run(["git", "-C", str(repo), "checkout", "-qb", "main"], check=True)
subprocess.run(["git", "-C", str(repo), "add", "-A"], check=True)
subprocess.run(
["git", "-C", str(repo), "-c", "user.email=t@t", "-c", "user.name=t",
"commit", "-qm", "init"], check=True,
)
subprocess.run(["git", "-C", str(repo), "checkout", "-qb", "feature"], check=True)
main_tf = repo / "main.tf"
main_tf.write_text(main_tf.read_text() + "\n# touched\n")
subprocess.run(["git", "-C", str(repo), "add", "-A"], check=True)
subprocess.run(
["git", "-C", str(repo), "-c", "user.email=t@t", "-c", "user.name=t",
"commit", "-qm", "touch"], check=True,
)
out_dir = tmp_path / "out"
plugin_cache = tmp_path / "plugin-cache"
plugin_cache.mkdir()
monkeypatch.setenv("TF_PLUGIN_CACHE_DIR", str(plugin_cache))
monkeypatch.setenv("TF_IN_AUTOMATION", "1")
mod = _load_cli()
rc = mod.main([
"--repo", str(repo),
"--base", "main",
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
manifest = json.loads((out_dir / "manifest.json").read_text())
assert manifest["plan_units"], (
f"no plan unit recorded; errors={manifest.get('errors')}"
)
pu = manifest["plan_units"][0]
assert pu["tool"] == "tofu"
assert pu["init"]["ok"] is True, pu["init"]
assert pu["plan"]["ok"] is True, pu["plan"]
assert rc == 0
@@ -0,0 +1,56 @@
from pathlib import Path
from scripts.reference_set import compute_consistency_norms
def _write(path: Path, body: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(body)
def test_norms_capture_repeated_attribute_presence(tmp_path):
_write(tmp_path / "live/prod/app/main.tf", "")
_write(tmp_path / "live/staging/app/main.tf", 'kms_key_arn = "a"\n')
_write(tmp_path / "live/dev/app/main.tf", 'kms_key_arn = "b"\n')
_write(tmp_path / "live/qa/app/main.tf", 'kms_key_arn = "c"\n')
norms = compute_consistency_norms(tmp_path, {
"live/prod/app": ["live/staging/app", "live/dev/app", "live/qa/app"],
})
assert norms["live/prod/app"]["attribute_norms"] == [
{
"attribute": "kms_key_arn",
"peer_dirs": ["live/dev/app", "live/qa/app", "live/staging/app"],
}
]
def test_norms_ignore_one_off_patterns(tmp_path):
_write(tmp_path / "envs/prod/main.tf", "")
_write(tmp_path / "envs/staging/main.tf", 'bucket_key_enabled = true\n')
_write(tmp_path / "envs/dev/main.tf", 'kms_key_arn = "b"\n')
norms = compute_consistency_norms(tmp_path, {
"envs/prod": ["envs/staging", "envs/dev"],
})
assert norms["envs/prod"]["attribute_norms"] == []
def test_norms_capture_repeated_module_wrappers(tmp_path):
_write(tmp_path / "modules/app-role/main.tf", "")
_write(tmp_path / "refs/one/main.tf", 'module "iam-policy-doc" {\n source = "../../modules/iam-policy-doc"\n}\n')
_write(tmp_path / "refs/two/main.tf", 'module "iam-policy-doc" {\n source = "../../modules/iam-policy-doc"\n}\n')
_write(tmp_path / "refs/three/main.tf", 'module "other" {\n source = "../../modules/other"\n}\n')
norms = compute_consistency_norms(tmp_path, {
"modules/app-role": ["refs/one", "refs/two", "refs/three"],
})
assert norms["modules/app-role"]["module_wrapper_norms"] == [
{
"module": "iam-policy-doc",
"peer_dirs": ["refs/one", "refs/two"],
}
]
@@ -0,0 +1,58 @@
import json
from datetime import datetime, timezone
from scripts.controls_schema import Control, ControlsFile
def test_controls_file_serializes_with_resource_type_index():
a = Control(
control_id="FSBP S3.5",
title="S3 buckets should require requests to use SSL",
severity="medium",
resource_types=["aws_s3_bucket"],
requirement="The bucket policy must include a deny statement for "
"non-TLS access (aws:SecureTransport = false).",
source_url="https://docs.aws.amazon.com/securityhub/.../S3.5",
)
b = Control(
control_id="FSBP IAM.5",
title="MFA should be enabled for IAM users",
severity="medium",
resource_types=["aws_iam_user"],
requirement="IAM users with console access must have MFA.",
source_url="https://docs.aws.amazon.com/securityhub/.../IAM.5",
)
cf = ControlsFile(
source="fsbp",
fetched_at=datetime(2026, 5, 12, 9, 0, 0, tzinfo=timezone.utc),
controls=[a, b],
)
payload = json.loads(cf.to_json())
assert payload["source"] == "fsbp"
assert payload["fetched_at"].endswith("+00:00") or payload["fetched_at"].endswith("Z")
by_rtype = payload["by_resource_type"]
assert "aws_s3_bucket" in by_rtype
assert by_rtype["aws_s3_bucket"] == ["FSBP S3.5"]
assert "aws_iam_user" in by_rtype
assert by_rtype["aws_iam_user"] == ["FSBP IAM.5"]
ids = {c["control_id"] for c in payload["controls"]}
assert ids == {"FSBP S3.5", "FSBP IAM.5"}
def test_control_appears_under_every_resource_type():
c = Control(
control_id="FSBP X.1",
title="t", severity="low",
resource_types=["aws_s3_bucket", "aws_s3_bucket_policy"],
requirement="r", source_url="u",
)
cf = ControlsFile(
source="fsbp",
fetched_at=datetime(2026, 5, 12, tzinfo=timezone.utc),
controls=[c],
)
payload = json.loads(cf.to_json())
assert payload["by_resource_type"]["aws_s3_bucket"] == ["FSBP X.1"]
assert payload["by_resource_type"]["aws_s3_bucket_policy"] == ["FSBP X.1"]
@@ -0,0 +1,71 @@
from pathlib import Path
from scripts.source_lookup import find_block
def _write(path: Path, body: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(body)
def test_review_context_collects_variables_and_locals(tmp_path):
_write(tmp_path / "main.tf", """
resource "aws_s3_bucket" "logs" {
kms_key_id = var.kms_key_id
acl = local.bucket_acl
}
""".strip() + "\n")
_write(tmp_path / "variables.tf", """
variable "kms_key_id" {
default = "alias/logs"
}
""".strip() + "\n")
_write(tmp_path / "locals.tf", """
locals {
bucket_acl = "private"
}
""".strip() + "\n")
block = find_block(tmp_path, "aws_s3_bucket.logs")
assert block is not None
assert block.review_context["variables"] == {"kms_key_id": "alias/logs"}
assert block.review_context["locals"] == {"bucket_acl": "private"}
def test_review_context_collects_related_policy_docs(tmp_path):
_write(tmp_path / "main.tf", """
data "aws_iam_policy_document" "bucket" {
statement {
actions = ["s3:GetObject"]
}
}
resource "aws_iam_policy" "bucket" {
policy = data.aws_iam_policy_document.bucket.json
}
""".strip() + "\n")
block = find_block(tmp_path, "aws_iam_policy.bucket")
assert block is not None
assert block.review_context["related_blocks"] == [
'data "aws_iam_policy_document" "bucket" {'
]
def test_review_context_collects_related_security_group_rules(tmp_path):
_write(tmp_path / "main.tf", """
resource "aws_security_group" "app" {
name = "app"
}
resource "aws_security_group_rule" "ingress_https" {
type = "ingress"
security_group_id = aws_security_group.app.id
}
""".strip() + "\n")
block = find_block(tmp_path, "aws_security_group.app")
assert block is not None
assert block.review_context["related_blocks"] == [
'resource "aws_security_group_rule" "ingress_https" {'
]
@@ -0,0 +1,63 @@
from unittest.mock import patch
from scripts.git_diff import changed_files, changed_dirs, changed_line_ranges
_DIFF = """\
diff --git a/live/prod/app/main.tf b/live/prod/app/main.tf
index abc..def 100644
--- a/live/prod/app/main.tf
+++ b/live/prod/app/main.tf
@@ -12 +12,2 @@
- acl = "old"
+ acl = "new"
+ versioning = true
diff --git a/modules/app-role/main.tf b/modules/app-role/main.tf
index 111..222 100644
--- a/modules/app-role/main.tf
+++ b/modules/app-role/main.tf
@@ -5,0 +6,2 @@
+ name = "x"
+ assume_role_policy = "..."
diff --git a/README.md b/README.md
index 333..444 100644
--- a/README.md
+++ b/README.md
@@ -1 +1 @@
-old
+new
"""
def _fake_run(args, **kwargs):
class R:
returncode = 0
stdout = _DIFF
stderr = ""
return R()
def test_changed_files_filters_tf_hcl():
with patch("subprocess.run", side_effect=_fake_run):
files = changed_files("/repo", "main", "HEAD")
assert "live/prod/app/main.tf" in files
assert "modules/app-role/main.tf" in files
assert "README.md" not in files
def test_changed_dirs_dedupes():
with patch("subprocess.run", side_effect=_fake_run):
dirs = changed_dirs("/repo", "main", "HEAD")
assert dirs == {"live/prod/app", "modules/app-role"}
def test_changed_line_ranges_parses_hunks():
with patch("subprocess.run", side_effect=_fake_run):
ranges = changed_line_ranges("/repo", "main", "HEAD",
"live/prod/app/main.tf")
assert ranges == [(12, 13)]
def test_changed_line_ranges_unknown_file():
with patch("subprocess.run", side_effect=_fake_run):
ranges = changed_line_ranges("/repo", "main", "HEAD", "no/such.tf")
assert ranges == []
@@ -0,0 +1,65 @@
import textwrap
from scripts.hcl_diff import find_resource_blocks, touched_resources
def test_finds_resource_block_spans(tmp_path):
body = textwrap.dedent("""\
variable "x" { default = 1 }
resource "aws_s3_bucket" "a" {
bucket = "x"
}
resource "aws_iam_role" "b" {
name = "y"
assume_role_policy = "..."
}
""")
f = tmp_path / "main.tf"
f.write_text(body)
blocks = find_resource_blocks(f)
by_name = {b.name: b for b in blocks}
assert by_name["a"].type == "aws_s3_bucket"
assert by_name["a"].start <= 3 <= by_name["a"].end
assert by_name["b"].type == "aws_iam_role"
assert by_name["b"].start <= 7 <= by_name["b"].end
def test_touched_resources_intersects(tmp_path):
body = textwrap.dedent("""\
resource "aws_s3_bucket" "a" {
bucket = "x"
}
resource "aws_iam_role" "b" {
name = "y"
}
""")
f = tmp_path / "main.tf"
f.write_text(body)
hits = touched_resources(f, added_ranges=[(6, 6)])
assert len(hits) == 1
assert hits[0].type == "aws_iam_role"
assert hits[0].name == "b"
def test_touched_resources_ignores_data_blocks(tmp_path):
body = textwrap.dedent("""\
data "aws_caller_identity" "current" {}
""")
f = tmp_path / "main.tf"
f.write_text(body)
hits = touched_resources(f, added_ranges=[(1, 1)])
assert hits == []
def test_touched_resources_empty_when_no_overlap(tmp_path):
body = textwrap.dedent("""\
resource "aws_s3_bucket" "a" {
bucket = "x"
}
""")
f = tmp_path / "main.tf"
f.write_text(body)
assert touched_resources(f, added_ranges=[(10, 12)]) == []
@@ -0,0 +1,93 @@
"""Tests for scripts/log-run.py."""
from __future__ import annotations
import importlib.util
import json
from pathlib import Path
_SPEC = importlib.util.spec_from_file_location(
"log_run",
Path(__file__).parent.parent / "scripts" / "log-run.py",
)
log_run = importlib.util.module_from_spec(_SPEC)
_SPEC.loader.exec_module(log_run)
def _write_usage(tmp_path: Path, usage: dict) -> Path:
p = tmp_path / "usage.json"
p.write_text(json.dumps(usage), encoding="utf-8")
return p
def _read_log(log: Path) -> list[dict]:
return [json.loads(line) for line in log.read_text(encoding="utf-8").splitlines() if line.strip()]
def test_writes_one_row_per_agent_with_finding_counts(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
(output / "findings-aws-bp-reviewer.json").write_text(
json.dumps({"findings": [{"resource": "r1", "control": "c1"}]}),
encoding="utf-8",
)
(output / "findings-walkthrough-reviewer.json").write_text(
json.dumps({"overview": "x", "plan_units": []}),
encoding="utf-8",
)
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"aws-bp-reviewer": {"model": "sonnet", "input_tokens": 100,
"output_tokens": 50, "duration_ms": 1234},
"walkthrough-reviewer": {"model": "sonnet", "input_tokens": 200,
"output_tokens": 80, "duration_ms": 4321},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "abc",
"--repo", "/tmp/repo", "--mode", "local",
"--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
rows = _read_log(log)
assert len(rows) == 2
by_agent = {r["agent"]: r for r in rows}
assert by_agent["aws-bp-reviewer"]["finding_count"] == 1
assert by_agent["walkthrough-reviewer"]["finding_count"] == 0
def test_missing_findings_file_counts_zero(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"phantom-reviewer": {"model": "sonnet", "input_tokens": 0,
"output_tokens": 0, "duration_ms": 0},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "ref", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
assert _read_log(log)[0]["finding_count"] == 0
def test_malformed_findings_file_counts_zero(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
(output / "findings-broken-reviewer.json").write_text("{bad", encoding="utf-8")
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"broken-reviewer": {"model": "haiku", "input_tokens": 10,
"output_tokens": 5, "duration_ms": 100},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "local", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
assert _read_log(log)[0]["finding_count"] == 0
@@ -0,0 +1,127 @@
import json
from scripts.manifest import (
Manifest, PlanUnit, PlanResult, InitResult, CatalogEntry, ModuleGraphEntry,
TrivyFinding,
)
def test_roundtrip_minimal():
m = Manifest(
base_ref="main",
head_ref="feat/x",
mode="local",
default_branch="main",
changed_source_dirs=[],
plan_units=[],
catalog=[],
trivy_findings=[],
module_graph={},
errors=[],
)
payload = json.loads(m.to_json())
assert payload["base_ref"] == "main"
assert payload["mode"] == "local"
assert payload["plan_units"] == []
def test_catalog_entry_required_fields():
from scripts.manifest import CatalogInstance
entry = CatalogEntry(
source_dir="live/prod/audit",
local_address="aws_s3_bucket.audit_logs",
type="aws_s3_bucket",
source="both",
instances=[CatalogInstance(
plan_dir="live/prod/audit",
address_at_plan="aws_s3_bucket.audit_logs",
action="create",
)],
block_header='resource "aws_s3_bucket" "audit_logs" {',
evidence_line='acl = "private"',
key_attributes={"acl": "private"},
review_context={"variables": {}, "locals": {}, "related_blocks": []},
block_file="main.tf",
block_start=1,
block_end=3,
)
d = entry.to_dict()
assert d["source"] == "both"
assert d["type"] == "aws_s3_bucket"
assert d["instances"][0]["plan_dir"] == "live/prod/audit"
assert d["block_header"].startswith('resource')
assert d["evidence_line"] == 'acl = "private"'
assert d["key_attributes"] == {"acl": "private"}
assert d["review_context"] == {"variables": {}, "locals": {}, "related_blocks": []}
def test_plan_unit_serializes_triggered_by():
pu = PlanUnit(
plan_dir="live/prod/app",
tool="terragrunt",
init=InitResult(ok=True, stdout_tail="ok", stderr_tail=""),
plan=PlanResult(ok=True, stdout_path="/tmp/x", exit_code=0,
summary="1 to add, 0 to change, 0 to destroy"),
triggered_by=["modules/app-role"],
terragrunt_changed=True,
changed_files=["live/prod/app/terragrunt.hcl"],
)
d = pu.to_dict()
assert d["tool"] == "terragrunt"
assert d["triggered_by"] == ["modules/app-role"]
assert d["init"]["ok"] is True
assert d["plan"]["exit_code"] == 0
assert d["terragrunt_changed"] is True
assert d["changed_files"] == ["live/prod/app/terragrunt.hcl"]
def test_trivy_findings_serialize():
finding = TrivyFinding(
check_id="AVD-AWS-0089",
title="S3 bucket allows public ACL",
severity="high",
message="Bucket ACL allows public access.",
file="main.tf",
start_line=21,
end_line=30,
resource_type="aws_s3_bucket",
)
assert finding.to_dict() == {
"check_id": "AVD-AWS-0089",
"title": "S3 bucket allows public ACL",
"severity": "high",
"message": "Bucket ACL allows public access.",
"file": "main.tf",
"start_line": 21,
"end_line": 30,
"resource_type": "aws_s3_bucket",
"source": "trivy",
}
def test_module_graph_serializes():
g = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app", "live/staging/app"],
sibling_modules_at_callsites=["modules/iam-policy-doc"],
)
}
findings = [
TrivyFinding(
check_id="AVD-AWS-0089",
title="S3 bucket allows public ACL",
severity="high",
message="Bucket ACL allows public access.",
file="main.tf",
resource_type="aws_s3_bucket",
)
]
m = Manifest(
base_ref="main", head_ref="x", mode="local", default_branch="main",
changed_source_dirs=[], plan_units=[], catalog=[], trivy_findings=findings,
module_graph=g, errors=[],
)
payload = json.loads(m.to_json())
assert payload["module_graph"]["modules/app-role"]["callsites"] == [
"live/prod/app", "live/staging/app",
]
assert payload["trivy_findings"][0]["check_id"] == "AVD-AWS-0089"
@@ -0,0 +1,33 @@
from scripts.manifest import Manifest, TflintFinding
def test_tflint_finding_to_dict():
f = TflintFinding(
rule="terraform_unused_declarations",
severity="warning",
message="variable \"foo\" is declared but not used",
file="modules/vpc/variables.tf",
start_line=12,
end_line=12,
link="https://github.com/terraform-linters/tflint-ruleset-terraform/blob/v0.14.1/docs/rules/terraform_unused_declarations.md",
)
d = f.to_dict()
assert d["rule"] == "terraform_unused_declarations"
assert d["source"] == "tflint"
def test_manifest_carries_tflint_findings():
m = Manifest(
base_ref="main", head_ref="HEAD", mode="local",
default_branch="main", changed_source_dirs=[],
plan_units=[], catalog=[], trivy_findings=[],
tflint_findings=[TflintFinding(
rule="terraform_required_version", severity="warning",
message="missing required_version", file="main.tf",
)],
module_graph={}, errors=[],
)
d = m.to_dict()
assert len(d["tflint_findings"]) == 1
assert d["tflint_findings"][0]["rule"] == "terraform_required_version"
assert d["tflint_findings"][0]["source"] == "tflint"
@@ -0,0 +1,84 @@
from pathlib import Path
from scripts.module_graph import build_module_graph
def _write(p: Path, body: str) -> None:
p.parent.mkdir(parents=True, exist_ok=True)
p.write_text(body)
def test_resolves_relative_source(tmp_path):
_write(tmp_path / "modules/app-role/main.tf",
'resource "aws_iam_role" "this" {}\n')
_write(tmp_path / "live/prod/app/main.tf", '''
module "app_role" {
source = "../../../modules/app-role"
}
''')
g = build_module_graph(tmp_path)
assert g["modules/app-role"].callsites == ["live/prod/app"]
def test_multiple_callsites_sorted(tmp_path):
_write(tmp_path / "modules/app-role/main.tf", "resource \"x\" \"y\" {}\n")
_write(tmp_path / "live/staging/app/main.tf",
'module "r" { source = "../../../modules/app-role" }\n')
_write(tmp_path / "live/prod/app/main.tf",
'module "r" { source = "../../../modules/app-role" }\n')
g = build_module_graph(tmp_path)
assert g["modules/app-role"].callsites == [
"live/prod/app", "live/staging/app",
]
def test_ignores_registry_and_git_sources(tmp_path):
_write(tmp_path / "live/prod/app/main.tf", '''
module "vpc" { source = "terraform-aws-modules/vpc/aws" }
module "other" { source = "git::https://example.com/x.git" }
''')
g = build_module_graph(tmp_path)
assert g == {}
def test_siblings_at_callsites_collected(tmp_path):
_write(tmp_path / "modules/app-role/main.tf", "resource \"x\" \"y\" {}\n")
_write(tmp_path / "modules/iam-policy-doc/main.tf",
"resource \"x\" \"y\" {}\n")
_write(tmp_path / "live/prod/app/main.tf", '''
module "role" { source = "../../../modules/app-role" }
module "policy" { source = "../../../modules/iam-policy-doc" }
''')
g = build_module_graph(tmp_path)
assert g["modules/app-role"].callsites == ["live/prod/app"]
assert "modules/iam-policy-doc" in g["modules/app-role"].sibling_modules_at_callsites
def test_callsite_local_names_track_sibling_modules(tmp_path):
_write(tmp_path / "modules/app-role/main.tf", 'resource "aws_iam_role" "this" {}\n')
_write(tmp_path / "modules/app-policy/main.tf",
'resource "aws_iam_policy" "this" {}\n')
_write(tmp_path / "live/prod/app/main.tf", '''
module "role" {
source = "../../../modules/app-role"
}
module "policy" {
source = "../../../modules/app-policy"
}
''')
g = build_module_graph(tmp_path)
assert g["modules/app-role"].callsite_local_names == {
"live/prod/app": {
"policy": "modules/app-policy",
"role": "modules/app-role",
}
}
assert g["modules/app-policy"].callsite_local_names == {
"live/prod/app": {
"policy": "modules/app-policy",
"role": "modules/app-role",
}
}
@@ -0,0 +1,39 @@
from pathlib import Path
from scripts.plan_output import parse_plan_resources
FIXTURE = Path(__file__).parent / "fixtures" / "plan_output_tofu.txt"
def test_parses_create_update_destroy_replace():
txt = FIXTURE.read_text()
resources = parse_plan_resources(txt)
by_addr = {r.address: r for r in resources}
assert by_addr["aws_s3_bucket.audit_logs"].action == "create"
assert by_addr["aws_s3_bucket.audit_logs"].type == "aws_s3_bucket"
assert by_addr["module.app_role.aws_iam_role.this"].action == "update"
assert by_addr["module.app_role.aws_iam_role.this"].type == "aws_iam_role"
assert by_addr["aws_security_group.legacy"].action == "delete"
assert by_addr["aws_iam_user.svc"].action == "replace"
def test_empty_when_no_changes():
txt = "No changes. Your infrastructure matches the configuration.\n"
assert parse_plan_resources(txt) == []
def test_type_for_data_source():
txt = """\
# data.aws_caller_identity.current will be read during apply
<= data "aws_caller_identity" "current" {
}
"""
out = parse_plan_resources(txt)
assert len(out) == 1
assert out[0].action == "read"
assert out[0].type == "aws_caller_identity"
assert out[0].address == "data.aws_caller_identity.current"
@@ -0,0 +1,67 @@
from unittest.mock import patch, MagicMock
from scripts.plan_runner import detect_tool, run_init, run_plan, extract_summary
def test_detect_tool_terragrunt(tmp_path):
(tmp_path / "terragrunt.hcl").write_text("# tg\n")
assert detect_tool(tmp_path) == "terragrunt"
def test_detect_tool_tofu(tmp_path):
(tmp_path / "main.tf").write_text("# tf\n")
assert detect_tool(tmp_path) == "tofu"
def test_extract_summary_finds_plan_line():
output = """\
Terraform will perform the following actions:
# aws_s3_bucket.x will be created
+ resource "aws_s3_bucket" "x" {
bucket = "y"
}
Plan: 1 to add, 0 to change, 0 to destroy.
"""
assert extract_summary(output) == "1 to add, 0 to change, 0 to destroy"
def test_extract_summary_no_changes():
output = "No changes. Your infrastructure matches the configuration.\n"
assert extract_summary(output) == "no changes"
def test_run_init_ok(tmp_path):
(tmp_path / "main.tf").write_text("# tf\n")
fake = MagicMock(returncode=0, stdout="Initializing...\nok\n", stderr="")
with patch("subprocess.run", return_value=fake) as p:
result = run_init(tmp_path, "tofu")
assert result.ok is True
assert "ok" in result.stdout_tail
p.assert_called_once()
args = p.call_args[0][0]
assert args[0] == "tofu"
assert "init" in args
def test_run_init_failure_records_stderr(tmp_path):
(tmp_path / "main.tf").write_text("# tf\n")
fake = MagicMock(returncode=1, stdout="", stderr="provider not found\n")
with patch("subprocess.run", return_value=fake):
result = run_init(tmp_path, "tofu")
assert result.ok is False
assert "provider not found" in result.stderr_tail
def test_run_plan_streams_to_file(tmp_path):
out_path = tmp_path / "plan.txt"
fake = MagicMock(returncode=0,
stdout="Plan: 2 to add, 0 to change, 0 to destroy.\n",
stderr="")
with patch("subprocess.run", return_value=fake):
result = run_plan(tmp_path, "tofu", out_path)
assert result.ok is True
assert result.exit_code == 0
assert result.summary == "2 to add, 0 to change, 0 to destroy"
assert out_path.read_text().endswith("0 to destroy.\n")
@@ -0,0 +1,49 @@
from pathlib import Path
import pytest
from scripts.plan_unit import classify_dir, DirKind
def _write(p: Path, body: str) -> None:
p.parent.mkdir(parents=True, exist_ok=True)
p.write_text(body)
def test_terragrunt_hcl_marks_plan_unit(tmp_path):
_write(tmp_path / "live/prod/app/terragrunt.hcl", 'include "root" { path = "x" }\n')
assert classify_dir(tmp_path, "live/prod/app") == DirKind.PLAN_UNIT
def test_tf_with_backend_block_marks_plan_unit(tmp_path):
body = '''
terraform {
required_version = ">= 1.5"
backend "s3" {
bucket = "x"
key = "y"
region = "us-east-1"
}
}
'''
_write(tmp_path / "live/prod/app/main.tf", body)
assert classify_dir(tmp_path, "live/prod/app") == DirKind.PLAN_UNIT
def test_tf_without_backend_marks_module(tmp_path):
body = '''
resource "aws_iam_role" "this" {
name = var.name
}
'''
_write(tmp_path / "modules/app-role/main.tf", body)
assert classify_dir(tmp_path, "modules/app-role") == DirKind.MODULE
def test_no_hcl_files_marks_unknown(tmp_path):
_write(tmp_path / "docs/readme.md", "hi")
assert classify_dir(tmp_path, "docs") == DirKind.UNKNOWN
def test_missing_dir_raises(tmp_path):
with pytest.raises(FileNotFoundError):
classify_dir(tmp_path, "nope")
@@ -0,0 +1,72 @@
from pathlib import Path
from scripts.reference_set import compute_consistency_norms, compute_reference_sets
from scripts.manifest import ModuleGraphEntry
def _touch(p: Path) -> None:
p.parent.mkdir(parents=True, exist_ok=True)
p.write_text("")
def test_module_change_uses_sibling_modules(tmp_path):
_touch(tmp_path / "modules/app-role/main.tf")
_touch(tmp_path / "modules/iam-policy-doc/main.tf")
_touch(tmp_path / "modules/eks-cluster/main.tf")
graph = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app"],
sibling_modules_at_callsites=["modules/eks-cluster",
"modules/iam-policy-doc"],
)
}
refs = compute_reference_sets(
tmp_path, changed_dirs={"modules/app-role"}, module_graph=graph,
)
assert refs["modules/app-role"] == [
"modules/eks-cluster", "modules/iam-policy-doc",
]
def test_terragrunt_component_uses_region_and_cross_env(tmp_path):
_touch(tmp_path / "live/prod/us-east-1/eks/terragrunt.hcl")
_touch(tmp_path / "live/prod/us-east-1/audit/terragrunt.hcl")
_touch(tmp_path / "live/prod/us-east-1/app/terragrunt.hcl")
_touch(tmp_path / "live/staging/us-east-1/eks/terragrunt.hcl")
_touch(tmp_path / "live/prod/us-west-2/eks/terragrunt.hcl")
refs = compute_reference_sets(
tmp_path, changed_dirs={"live/prod/us-east-1/eks"}, module_graph={},
)
ref = set(refs["live/prod/us-east-1/eks"])
assert "live/prod/us-east-1/audit" in ref
assert "live/prod/us-east-1/app" in ref
assert "live/staging/us-east-1/eks" in ref
assert "live/prod/us-west-2/eks" in ref
assert "live/prod/us-east-1/eks" not in ref
def test_plain_terraform_uses_sibling_dirs(tmp_path):
_touch(tmp_path / "envs/prod/main.tf")
_touch(tmp_path / "envs/staging/main.tf")
_touch(tmp_path / "envs/dev/main.tf")
refs = compute_reference_sets(
tmp_path, changed_dirs={"envs/prod"}, module_graph={},
)
assert set(refs["envs/prod"]) == {"envs/staging", "envs/dev"}
def test_compute_consistency_norms_ignores_single_peer_drift(tmp_path):
(tmp_path / "envs/prod").mkdir(parents=True, exist_ok=True)
(tmp_path / "envs/staging").mkdir(parents=True, exist_ok=True)
(tmp_path / "envs/dev").mkdir(parents=True, exist_ok=True)
(tmp_path / "envs/qa").mkdir(parents=True, exist_ok=True)
(tmp_path / "envs/staging/main.tf").write_text('kms_key_arn = "a"\n')
(tmp_path / "envs/dev/main.tf").write_text('kms_key_arn = "b"\n')
(tmp_path / "envs/qa/main.tf").write_text('bucket_key_enabled = true\n')
refs = {"envs/prod": ["envs/staging", "envs/dev", "envs/qa"]}
norms = compute_consistency_norms(tmp_path, refs)
attr_norms = norms["envs/prod"]["attribute_norms"]
assert {"attribute": "kms_key_arn", "peer_dirs": ["envs/dev", "envs/staging"]} in attr_norms
assert not any(norm["attribute"] == "bucket_key_enabled" for norm in attr_norms)
@@ -0,0 +1,59 @@
import importlib.util
import json
from pathlib import Path
from unittest.mock import patch
_SKILL_ROOT = Path(__file__).resolve().parent.parent
def _load_cli():
spec = importlib.util.spec_from_file_location(
"refresh_controls", _SKILL_ROOT / "scripts" / "refresh-controls.py"
)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
_FSBP_HTML = (_SKILL_ROOT / "tests" / "fixtures" / "fsbp_index_sample.html").read_text()
_CIS_HTML = (_SKILL_ROOT / "tests" / "fixtures" / "cis_sample.html").read_text()
class _FakeResponse:
def __init__(self, text: str, status_code: int = 200):
self.text = text
self.status_code = status_code
def raise_for_status(self):
if self.status_code >= 400:
raise RuntimeError(f"HTTP {self.status_code}")
def _fake_get(url, **kw):
if "fsbp" in url or "foundational-security-best-practices" in url:
return _FakeResponse(_FSBP_HTML)
if "cis" in url:
return _FakeResponse(_CIS_HTML)
return _FakeResponse("<html></html>")
def test_refresh_writes_three_json_files(tmp_path):
mod = _load_cli()
out_dir = tmp_path / "data" / "controls"
with patch("requests.get", side_effect=_fake_get):
rc = mod.main(["--output-dir", str(out_dir)])
assert rc == 0
fsbp = json.loads((out_dir / "fsbp.json").read_text())
cis = json.loads((out_dir / "cis.json").read_text())
meta = json.loads((out_dir / "meta.json").read_text())
assert fsbp["source"] == "fsbp"
assert "aws_s3_bucket" in fsbp["by_resource_type"]
assert any(c["control_id"] == "FSBP S3.5" for c in fsbp["controls"])
assert cis["source"] == "cis"
assert any(c["control_id"] == "CIS Account.1" for c in cis["controls"])
assert meta["fsbp"]["url"]
assert meta["fsbp"]["fetched_at"]
assert meta["cis"]["url"]
@@ -0,0 +1,92 @@
from pathlib import Path
from scripts.resolve_plan_units import resolve_plan_units
from scripts.manifest import ModuleGraphEntry
def _mk(tmp_path: Path, rel: str, has_backend: bool = False, terragrunt: bool = False):
d = tmp_path / rel
d.mkdir(parents=True, exist_ok=True)
if terragrunt:
(d / "terragrunt.hcl").write_text("# tg\n")
if has_backend:
(d / "main.tf").write_text(
'terraform { backend "local" {} }\n'
'resource "null_resource" "x" {}\n'
)
def test_plan_unit_dir_planned_directly(tmp_path):
_mk(tmp_path, "live/prod/app", has_backend=True)
plan_units, orphans = resolve_plan_units(
repo_root=tmp_path,
changed_dirs={"live/prod/app"},
module_graph={},
)
assert plan_units == {"live/prod/app": ["live/prod/app"]}
assert orphans == []
def test_terragrunt_dir_planned_directly(tmp_path):
_mk(tmp_path, "live/prod/app", terragrunt=True)
plan_units, orphans = resolve_plan_units(
repo_root=tmp_path,
changed_dirs={"live/prod/app"},
module_graph={},
)
assert plan_units == {"live/prod/app": ["live/prod/app"]}
assert orphans == []
def test_module_change_expands_to_callsites(tmp_path):
_mk(tmp_path, "live/prod/app", has_backend=True)
_mk(tmp_path, "live/staging/app", has_backend=True)
_mk(tmp_path, "modules/app-role")
(tmp_path / "modules/app-role/main.tf").write_text(
'resource "null_resource" "x" {}\n')
graph = {
"modules/app-role": ModuleGraphEntry(
callsites=["live/prod/app", "live/staging/app"],
)
}
plan_units, orphans = resolve_plan_units(
repo_root=tmp_path,
changed_dirs={"modules/app-role"},
module_graph=graph,
)
assert plan_units == {
"live/prod/app": ["modules/app-role"],
"live/staging/app": ["modules/app-role"],
}
assert orphans == []
def test_module_with_no_callsites_is_orphan(tmp_path):
_mk(tmp_path, "modules/lonely")
(tmp_path / "modules/lonely/main.tf").write_text(
'resource "null_resource" "x" {}\n')
plan_units, orphans = resolve_plan_units(
repo_root=tmp_path,
changed_dirs={"modules/lonely"},
module_graph={},
)
assert plan_units == {}
assert orphans == ["modules/lonely"]
def test_triggered_by_merges_when_both_dir_and_module_change(tmp_path):
_mk(tmp_path, "live/prod/app", has_backend=True)
_mk(tmp_path, "modules/app-role")
(tmp_path / "modules/app-role/main.tf").write_text(
'resource "null_resource" "x" {}\n')
graph = {
"modules/app-role": ModuleGraphEntry(callsites=["live/prod/app"]),
}
plan_units, _ = resolve_plan_units(
repo_root=tmp_path,
changed_dirs={"live/prod/app", "modules/app-role"},
module_graph=graph,
)
assert plan_units == {
"live/prod/app": ["live/prod/app", "modules/app-role"],
}
@@ -0,0 +1,50 @@
import pytest
from pathlib import Path
from scripts.review_stats import compute_stats
from scripts.telemetry import append_subagent_run, append_verdict
def _seed(log: Path):
append_subagent_run(
log, run_id="r1", repo="/r", mode="local",
agent="aws-bp-reviewer", model="claude-sonnet-4-6",
input_tokens=1000, output_tokens=200, duration_ms=5000, finding_count=3,
)
append_verdict(log, run_id="r1", agent="aws-bp-reviewer",
rule_id="TRIVY AVD-AWS-0089", file="a.tf", line=1, verdict="kept")
append_verdict(log, run_id="r1", agent="aws-bp-reviewer",
rule_id="TRIVY AVD-AWS-0089", file="b.tf", line=2, verdict="dismissed")
append_verdict(log, run_id="r1", agent="aws-bp-reviewer",
rule_id="TRIVY AVD-AWS-0089", file="c.tf", line=3, verdict="false_positive")
def test_compute_stats_precision_and_tokens(tmp_path):
log = tmp_path / "runs.jsonl"
_seed(log)
stats = compute_stats(log)
assert stats["runs"] == 1
a = stats["by_agent"]["aws-bp-reviewer"]
assert a["total"] == 3
assert a["kept"] == 1
assert a["dismissed"] == 1
assert a["false_positive"] == 1
assert a["precision"] == pytest.approx(1 / 3)
assert a["tokens"] == 1200
assert a["tokens_per_kept"] == 1200
def test_compute_stats_empty(tmp_path):
stats = compute_stats(tmp_path / "missing.jsonl")
assert stats == {"by_agent": {}, "by_rule": {}, "runs": 0}
def test_compute_stats_per_rule(tmp_path):
log = tmp_path / "runs.jsonl"
_seed(log)
stats = compute_stats(log)
rule_key = "aws-bp-reviewer/TRIVY AVD-AWS-0089"
r = stats["by_rule"][rule_key]
assert r["total"] == 3
assert r["kept"] == 1
assert r["precision"] == pytest.approx(1 / 3)
@@ -0,0 +1,109 @@
from scripts.scanners import (
_normalize_tflint_findings,
_run_tflint,
)
def test_normalize_tflint_findings_parses_issues():
payload = {
"issues": [
{
"rule": {
"name": "terraform_unused_declarations",
"severity": "warning",
"link": "https://example.com/rule",
},
"message": "variable \"foo\" is declared but not used",
"range": {
"filename": "variables.tf",
"start": {"line": 12},
"end": {"line": 12},
},
}
],
"errors": [],
}
findings = _normalize_tflint_findings(payload, scanned_dir="modules/vpc")
assert len(findings) == 1
f = findings[0]
assert f.rule == "terraform_unused_declarations"
assert f.file == "modules/vpc/variables.tf"
assert f.start_line == 12
assert f.severity == "warning"
assert f.source == "tflint"
def test_normalize_tflint_findings_preserves_absolute_filename():
payload = {
"issues": [
{
"rule": {"name": "r", "severity": "warning"},
"message": "m",
"range": {"filename": "modules/vpc/variables.tf",
"start": {"line": 1}, "end": {"line": 1}},
}
]
}
findings = _normalize_tflint_findings(payload, scanned_dir="modules/vpc")
assert findings[0].file == "modules/vpc/variables.tf"
def test_normalize_tflint_findings_empty():
assert _normalize_tflint_findings({"issues": []}, scanned_dir=".") == []
def test_run_tflint_handles_missing_binary(tmp_path, monkeypatch):
monkeypatch.setenv("PATH", "/nonexistent")
payload, findings, err = _run_tflint(tmp_path, ["main.tf"])
assert findings == []
assert payload == {}
assert err is None
def test_run_tflint_skips_when_no_terraform_files(tmp_path):
payload, findings, err = _run_tflint(tmp_path, ["README.md", "Makefile"])
assert findings == []
assert err is None
def test_normalize_tflint_findings_collapses_dot_scanned_dir():
payload = {
"issues": [
{
"rule": {"name": "r", "severity": "warning"},
"message": "m",
"range": {"filename": "main.tf",
"start": {"line": 1}, "end": {"line": 1}},
}
]
}
findings = _normalize_tflint_findings(payload, scanned_dir=".")
assert findings[0].file == "main.tf"
def test_run_tflint_aggregates_errors_from_multiple_dirs(tmp_path, monkeypatch):
(tmp_path / "modules" / "a").mkdir(parents=True)
(tmp_path / "modules" / "b").mkdir(parents=True)
(tmp_path / "modules" / "a" / "main.tf").write_text("")
(tmp_path / "modules" / "b" / "main.tf").write_text("")
from scripts import scanners
def fake_run(cmd, **_):
class R:
returncode = 1
stdout = ""
stderr = "boom"
return R()
monkeypatch.setattr(scanners.subprocess, "run", fake_run)
monkeypatch.setattr(scanners.shutil, "which", lambda _: "/usr/bin/tflint")
payload, findings, err = scanners._run_tflint(
tmp_path, ["modules/a/main.tf", "modules/b/main.tf"]
)
assert findings == []
assert err is not None
assert "modules/a" in err
assert "modules/b" in err
@@ -0,0 +1,37 @@
from pathlib import Path
from scripts.scrape_cis import parse_cis_page
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_cis_unified_controls():
html = (_FIXTURES / "cis_sample.html").read_text()
rows = parse_cis_page(html, source_url="https://docs.aws.amazon.com/.../cis.html")
by_id = {r.control_id: r for r in rows}
assert "CIS Account.1" in by_id
a = by_id["CIS Account.1"]
assert "Security contact" in a.title
assert "v5.0.0 §1.2" in a.requirement
assert "v1.2.0 §1.18" in a.requirement
assert a.severity == "medium"
def test_skips_empty_version_cells():
html = (_FIXTURES / "cis_sample.html").read_text()
rows = parse_cis_page(html, source_url="https://docs.aws.amazon.com/.../cis.html")
by_id = {r.control_id: r for r in rows}
iam5 = by_id["CIS IAM.5"]
assert "v1.4.0" not in iam5.requirement
assert "v5.0.0 §1.10" in iam5.requirement
assert "v1.2.0 §1.2" in iam5.requirement
def test_resource_types_via_prefix_map():
html = (_FIXTURES / "cis_sample.html").read_text()
rows = parse_cis_page(html, source_url="https://docs.aws.amazon.com/.../cis.html")
by_id = {r.control_id: r for r in rows}
assert "aws_s3_bucket" in by_id["CIS S3.5"].resource_types
assert "aws_iam_user" in by_id["CIS IAM.5"].resource_types
assert "aws_cloudtrail" in by_id["CIS CloudTrail.1"].resource_types
@@ -0,0 +1,38 @@
from pathlib import Path
from scripts.scrape_fsbp import parse_fsbp_index
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_fsbp_index_rows():
html = (_FIXTURES / "fsbp_index_sample.html").read_text()
rows = parse_fsbp_index(html, base_url="https://docs.aws.amazon.com")
by_id = {r.control_id: r for r in rows}
assert "S3.5" in by_id
assert by_id["S3.5"].title == "S3 general purpose buckets should require requests to use SSL"
assert by_id["S3.5"].severity == "medium"
assert by_id["S3.5"].detail_url.endswith("s3-controls.html#s3-5")
assert "IAM.5" in by_id
assert "Account.1" in by_id
def test_resource_types_attached_via_prefix_map():
html = (_FIXTURES / "fsbp_index_sample.html").read_text()
rows = parse_fsbp_index(html, base_url="https://docs.aws.amazon.com")
by_id = {r.control_id: r for r in rows}
assert "aws_s3_bucket" in by_id["S3.5"].resource_types
assert "aws_iam_user" in by_id["IAM.5"].resource_types
assert "aws_account_alternate_contact" in by_id["Account.1"].resource_types
def test_dedups_repeat_control_ids():
"""A control linked multiple times on the page should appear once."""
html = """\
<html><body>
<p><a href="./s3-controls.html#s3-1">[S3.1] First link</a></p>
<p><a href="./s3-controls.html#s3-1">[S3.1] Second link to same control</a></p>
</body></html>"""
rows = parse_fsbp_index(html, base_url="https://docs.aws.amazon.com")
assert len([r for r in rows if r.control_id == "S3.1"]) == 1
@@ -0,0 +1,52 @@
from scripts.slicing import _slice_for_agent
def _sample_manifest():
return {
"base_ref": "main", "head_ref": "HEAD", "mode": "local",
"default_branch": "main", "errors": [], "changed_source_dirs": [],
"catalog": [{"type": "aws_s3_bucket"}],
"plan_units": [],
"module_graph": {},
"trivy_findings": [{"rule": "AVD-1"}],
"tflint_findings": [{"rule": "terraform_required_version"}],
}
def test_aws_bp_slice_includes_tflint():
sliced = _slice_for_agent(_sample_manifest(), "aws-bp")
assert sliced["tflint_findings"][0]["rule"] == "terraform_required_version"
def test_fsbp_slice_includes_tflint():
sliced = _slice_for_agent(_sample_manifest(), "fsbp")
assert sliced["tflint_findings"][0]["rule"] == "terraform_required_version"
def test_cis_slice_includes_tflint():
sliced = _slice_for_agent(_sample_manifest(), "cis")
assert sliced["tflint_findings"][0]["rule"] == "terraform_required_version"
def test_aws_bp_slice_still_includes_trivy():
sliced = _slice_for_agent(_sample_manifest(), "aws-bp")
assert sliced["trivy_findings"][0]["rule"] == "AVD-1"
def test_consistency_slice_includes_tflint():
sliced = _slice_for_agent(_sample_manifest(), "consistency")
assert sliced["tflint_findings"][0]["rule"] == "terraform_required_version"
def test_hygiene_slice_includes_tflint_and_omits_trivy():
sliced = _slice_for_agent(_sample_manifest(), "tf-hygiene")
assert sliced["tflint_findings"][0]["rule"] == "terraform_required_version"
assert "trivy_findings" not in sliced
assert "plan_units" in sliced
assert "catalog" in sliced
assert "changed_source_dirs" in sliced
def test_unknown_agent_returns_full_manifest():
sliced = _slice_for_agent(_sample_manifest(), "unknown-agent")
assert sliced == _sample_manifest()
@@ -0,0 +1,22 @@
from pathlib import Path
SKILL_DIR = Path(__file__).resolve().parent.parent
def test_imports():
from scripts import manifest # noqa: F401
def test_skill_documents_trivy_first_flow():
skill = (SKILL_DIR / "SKILL.md").read_text()
aws_bp = (SKILL_DIR / "agents/aws-bp-reviewer.md").read_text()
consistency = (SKILL_DIR / "agents/consistency-reviewer.md").read_text()
design = (SKILL_DIR / "README.md").read_text()
assert "trivy-findings.json" in skill
assert "Agents: `walkthrough-reviewer`, `aws-bp-reviewer`, `consistency-reviewer`,\n`tf-hygiene-reviewer`." in skill
assert "Trivy as the first-pass scanner" in aws_bp
assert "`trivy config` first-pass scan" in design
assert "CONSISTENCY_NORMS" in consistency
assert "consistency_norms.json" in design
assert "block_text" not in design
@@ -0,0 +1,57 @@
import textwrap
from scripts.source_lookup import find_block
def test_finds_block_in_single_file(tmp_path):
main_tf = tmp_path / "main.tf"
main_tf.write_text(textwrap.dedent("""\
resource "aws_s3_bucket" "logs" {
bucket = "logs"
}
resource "aws_iam_role" "svc" {
name = "svc"
}
"""))
loc = find_block(tmp_path, "aws_iam_role.svc")
assert loc is not None
assert loc.file.name == "main.tf"
assert loc.start_line == 5
assert loc.end_line == 7
assert 'resource "aws_iam_role" "svc"' in loc.text
assert loc.text.rstrip().endswith("}")
def test_searches_multiple_tf_files(tmp_path):
(tmp_path / "a.tf").write_text(
'resource "aws_s3_bucket" "a" { bucket = "a" }\n'
)
(tmp_path / "b.tf").write_text(
'resource "aws_s3_bucket" "b" { bucket = "b" }\n'
)
loc = find_block(tmp_path, "aws_s3_bucket.b")
assert loc is not None
assert loc.file.name == "b.tf"
assert 'bucket = "b"' in loc.text
def test_returns_none_when_missing(tmp_path):
(tmp_path / "main.tf").write_text(
'resource "aws_s3_bucket" "logs" { bucket = "x" }\n'
)
assert find_block(tmp_path, "aws_iam_role.svc") is None
def test_ignores_data_blocks(tmp_path):
(tmp_path / "main.tf").write_text(textwrap.dedent("""\
data "aws_caller_identity" "current" {}
resource "aws_iam_role" "real" {
name = "x"
}
"""))
loc = find_block(tmp_path, "aws_caller_identity.current")
assert loc is None
loc2 = find_block(tmp_path, "aws_iam_role.real")
assert loc2 is not None
@@ -0,0 +1,43 @@
import pytest
from scripts.telemetry import append_subagent_run, append_verdict, read_runs
def test_round_trip(tmp_path):
log = tmp_path / "runs.jsonl"
append_subagent_run(
log, run_id="r1", repo="/r", mode="local",
agent="aws-bp-reviewer", model="claude-sonnet-4-6",
input_tokens=100, output_tokens=20, duration_ms=1000, finding_count=3,
)
append_verdict(
log, run_id="r1", agent="aws-bp-reviewer",
rule_id="TRIVY AVD-AWS-0089", file="vpc.tf", line=12,
verdict="kept", notes="real",
)
records = read_runs(log)
assert len(records) == 2
assert records[0]["kind"] == "subagent_run"
assert records[0]["input_tokens"] == 100
assert records[1]["kind"] == "verdict"
assert records[1]["verdict"] == "kept"
def test_append_creates_parent_dir(tmp_path):
log = tmp_path / "deep" / "nested" / "runs.jsonl"
append_subagent_run(
log, run_id="r2", repo="/r", mode="local",
agent="x", model="y", input_tokens=0, output_tokens=0,
duration_ms=0, finding_count=0,
)
assert log.exists()
def test_invalid_verdict_raises(tmp_path):
log = tmp_path / "runs.jsonl"
with pytest.raises(ValueError, match="invalid verdict"):
append_verdict(
log, run_id="r3", agent="aws-bp-reviewer",
rule_id="TRIVY AVD-AWS-0089", file="a.tf", line=5,
verdict="maybe", notes="",
)
+3
View File
@@ -0,0 +1,3 @@
version = 1
revision = 3
requires-python = ">=3.12"