Move the code and terraform audits into the reviews plugin
Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.
Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.
Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.
Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.
Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.
Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.
Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
@@ -0,0 +1,287 @@
|
||||
# audit-code
|
||||
|
||||
Automated, linter-driven audit of source-code changes. Runs an OSS SAST and
|
||||
quality stack against a repo, filters every finding down to lines the change
|
||||
actually touched, slices the result per reviewer concern, and fans out eight
|
||||
subagents in parallel to triage.
|
||||
|
||||
Part of the `reviews` plugin (marketplace `mroberts`). The sibling skill
|
||||
`review-pr` does the opposite job — a guided walkthrough for a human reading a
|
||||
PR. `audit-code` is the tool-driven one; it produces findings, not a reading
|
||||
plan.
|
||||
|
||||
Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell,
|
||||
GitHub Actions workflows.
|
||||
|
||||
## When it runs
|
||||
|
||||
Auto-activates on phrasing like "audit this code", "run the linters on this
|
||||
change", or an explicit `/audit-code`.
|
||||
|
||||
Two modes:
|
||||
|
||||
| Invocation | Mode | Diff target | Output |
|
||||
|---------------------------|---------|-------------------------|---------------------------------|
|
||||
| `/audit-code` | `local` | working tree vs origin | interactive walkthrough in chat |
|
||||
| `/audit-code <ref>` | `ref` | `<ref>` vs origin | `audit-code-<short>.md` + chat |
|
||||
| `/audit-code <pr_number>` | `ref` | PR head ref vs base | same as ref mode |
|
||||
|
||||
Local mode runs in the user's current checkout. Ref/PR mode isolates the
|
||||
checkout first — a git worktree when the cwd is a git checkout of the target
|
||||
repo, otherwise a fresh `gh repo clone` into
|
||||
`~/.claude/cache/audit-code/<short-ref>/`. The clone path is what makes the
|
||||
skill work from jj workspaces and from outside the repo entirely.
|
||||
|
||||
PR-number resolution requires `gh` on PATH and authenticated. The repo is
|
||||
resolved to `owner/repo` up front and passed as `--repo` on every `gh` call,
|
||||
so `gh` never autodetects from cwd.
|
||||
|
||||
## Pipeline
|
||||
|
||||
```
|
||||
SKILL.md (orchestrator)
|
||||
│
|
||||
├─► scripts/collect-findings.py --repo --head --output-dir --mode
|
||||
│ ├── resolve default branch, fetch origin, compute base
|
||||
│ ├── git diff --unified=0 base...head → changed files + added-line ranges
|
||||
│ ├── classify each path: supported source / GHA / dep manifest / skipped
|
||||
│ ├── dispatch linters for the languages actually present
|
||||
│ ├── diff-filter: drop findings not overlapping an added-line range
|
||||
│ ├── write manifest.json
|
||||
│ └── write 8 per-agent manifest slices
|
||||
│
|
||||
├─► 8 Task subagents IN PARALLEL, one per slice
|
||||
│ each writes findings-<agent>.json
|
||||
│
|
||||
├─► aggregate: dedupe on {file, line, rule_id}, group by severity
|
||||
├─► scripts/log-run.py → append telemetry rows
|
||||
└─► chat summary (local) or markdown report (ref)
|
||||
```
|
||||
|
||||
`collect-findings.py` exits non-zero when the diff cannot be computed, or when
|
||||
the change set contains no supported source files and no dependency manifests.
|
||||
On failure it still writes `manifest.json`; the `errors[]` array holds the
|
||||
reason.
|
||||
|
||||
Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30
|
||||
others — see `scripts/language_detect.py`) do not abort the run. Supported
|
||||
files are still reviewed, and one `errors[]` entry per language records the
|
||||
gap so partial coverage is visible rather than silent.
|
||||
|
||||
### Diff filtering
|
||||
|
||||
`scripts/diff_filter.py` keeps a finding only if its `[line, end_line]` span
|
||||
overlaps one of the added-line ranges recorded for that file. Paths are
|
||||
normalized (absolute → repo-relative, `./` stripped) before comparison, since
|
||||
adapters emit paths in whichever form their tool produced. Findings in
|
||||
untouched code are discarded, so the audit is scoped to the change.
|
||||
|
||||
### Tool resolution
|
||||
|
||||
`scripts/runner.py` resolves each binary against `.venv/bin/` inside the skill
|
||||
directory first, then falls back to `PATH`. A tool that resolves nowhere is
|
||||
recorded as `ran: false` with reason `not on PATH` and the run continues —
|
||||
missing tools degrade coverage, they do not fail the audit. Every tool
|
||||
invocation has a 180-second timeout.
|
||||
|
||||
## External tools
|
||||
|
||||
All of these are shelled out to. Nothing is vendored.
|
||||
|
||||
| Tool | Fires when | Purpose |
|
||||
|---|---|---|
|
||||
| `bandit` | Python in diff | Python security |
|
||||
| `ruff` | Python in diff | Python lint (incl. S-rules) |
|
||||
| `ruff` (idiom pass) | Python in diff | second `ruff` run, `--isolated --select SIM,PERF,UP,RET,PLR,C90,B`; reported as tool `ruff-idiom` |
|
||||
| `mypy` | Python in diff | Python types |
|
||||
| `eslint` | JS/TS in diff | JS/TS lint + security plugins |
|
||||
| `tsc` | JS/TS in diff | TypeScript types (`--noEmit`) |
|
||||
| `dotnet` | C# in diff | `dotnet build --no-incremental`; surfaces Roslyn + SecurityCodeScan (`SCS*`) |
|
||||
| `selene` | Lua in diff | Lua lint |
|
||||
| `luac` | Lua in diff | `luac -p` per file; syntax/parse errors |
|
||||
| `pwsh` + `PSScriptAnalyzer` | PowerShell in diff | PowerShell lint; security rules split out |
|
||||
| `pwsh` + `InjectionHunter` | PowerShell in diff | PowerShell injection SAST |
|
||||
| `actionlint` | `.github/workflows/` or `.github/actions/` YAML in diff | workflow correctness |
|
||||
| `zizmor` | same as above | workflow security (SARIF output) |
|
||||
| `opengrep` | always | polyglot SAST, `--config=auto` plus bundled `scripts/rules/lua-security.yaml` |
|
||||
| `gitleaks` | always | secret detection |
|
||||
| `pip-audit` | always | Python dependency CVEs |
|
||||
| `osv-scanner` | always | multi-ecosystem dependency CVEs |
|
||||
| `lizard` | always | multi-language cyclomatic complexity |
|
||||
| `vulture` | Python in diff | dead code |
|
||||
| `radon` | Python in diff | complexity |
|
||||
| `interrogate` | Python in diff | docstring coverage |
|
||||
| `knip` (via `npx`) | JS/TS in diff | unused exports/files |
|
||||
| `jscpd` | JS/TS in diff | copy-paste duplication |
|
||||
|
||||
Dependency manifests (`requirements*.txt`, `package.json`, lockfiles,
|
||||
`*.csproj`, `*.sln`, `pyproject.toml`, …) are detected separately from source
|
||||
files. `requirements*.txt` and `package.json` additionally get a structured
|
||||
before/after package diff (added / removed / upgraded) computed by
|
||||
`scripts/package_diff.py`.
|
||||
|
||||
PowerShell is worth calling out: if `pwsh` is missing, both
|
||||
`psscriptanalyzer` and `injectionhunter` report unavailable. InjectionHunter
|
||||
is the only injection SAST pass for PowerShell, so its absence means
|
||||
PowerShell injection flaws went unchecked — that is a coverage gap, not a
|
||||
clean result.
|
||||
|
||||
### Installing them
|
||||
|
||||
```
|
||||
scripts/install-tools.sh
|
||||
```
|
||||
|
||||
What it does:
|
||||
|
||||
1. `uv sync --group tools` into `${SKILL_DIR}/.venv/` — installs `bandit`,
|
||||
`ruff`, `mypy`, `pip-audit`, `vulture`, `radon`, `interrogate`, `lizard`.
|
||||
Falls back to `pip install bandit ruff mypy pip-audit` if `uv` is absent.
|
||||
2. Downloads a platform-matched `opengrep` release binary into
|
||||
`${SKILL_DIR}/.venv/bin/` (Linux x86_64/aarch64, macOS x86_64/arm64).
|
||||
3. Installs the `PSScriptAnalyzer` and `InjectionHunter` PowerShell modules
|
||||
for the current user, if `pwsh` is present.
|
||||
4. Installs `gitleaks`, `osv-scanner`, and `gh` via Homebrew, paru/yay/pacman,
|
||||
or prints apt-era install instructions.
|
||||
5. Prints a resolution table for each tool, then lists the per-project tools
|
||||
it deliberately does not install.
|
||||
|
||||
Not installed by the script — these must exist in the target repo or on PATH
|
||||
yourself: `eslint` (+ `eslint-plugin-security`), `typescript`/`tsc`, `dotnet`
|
||||
with `SecurityCodeScan`, `knip`, `jscpd`, `selene`, `luac`, `actionlint`,
|
||||
`zizmor`.
|
||||
|
||||
## The eight subagents
|
||||
|
||||
Each subagent gets the prompt from `agents/<name>.md` plus four substituted
|
||||
variables: `MANIFEST`, `REPO`, `MODE`, `OUTPUT`. Each writes
|
||||
`findings-<agent>.json` into the output directory.
|
||||
|
||||
| Agent | Manifest slice | Covers |
|
||||
|---|---|---|
|
||||
| `walkthrough-reviewer` | `manifest-walkthrough.json` | Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. |
|
||||
| `security-triage-reviewer` | `manifest-security-triage.json` | bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, `SCS*` from dotnet, and the PSScriptAnalyzer credential/crypto rules. |
|
||||
| `type-safety-reviewer` | `manifest-type-safety.json` | mypy, tsc, non-security eslint, non-`SCS` dotnet. Also hunts leaked `Any`, unexplained `type: ignore`/`@ts-ignore`, missing hints on new public functions. |
|
||||
| `dependency-reviewer` | `manifest-dependency.json` | pip-audit and osv-scanner findings, plus the `package_diffs` added/removed/upgraded entries with advisory lookups. |
|
||||
| `consistency-reviewer` | `manifest-consistency.json` | Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the `ruff-idiom` findings. |
|
||||
| `secrets-reviewer` | `manifest-secrets.json` | gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. |
|
||||
| `maintainability-reviewer` | `manifest-maintainability.json` | vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the `C901`/`PLR0915` complexity rules from the ruff idiom pass. |
|
||||
| `gha-reviewer` | `manifest-gha-reviewer.json` | actionlint and zizmor findings against changed workflow files. |
|
||||
|
||||
Slicing logic lives in `scripts/slicing.py`, which is the authority on which
|
||||
tool's output reaches which agent.
|
||||
|
||||
Default models: Sonnet for seven agents, Haiku 4.5 for
|
||||
`maintainability-reviewer` (its work is deterministic-tool triage). Each has
|
||||
an override flag — `--security-model`, `--type-safety-model`,
|
||||
`--dependency-model`, `--consistency-model`, `--secrets-model`,
|
||||
`--maintainability-model`, `--walkthrough-model`, `--gha-model`.
|
||||
|
||||
Aggregation dedupes findings sharing `{file, line, rule_id}` — first to finish
|
||||
wins, and the loser is recorded under `also_flagged_by`. Severities are
|
||||
`critical`, `high`, `medium`, `low`, `info`.
|
||||
|
||||
## Output artifacts
|
||||
|
||||
Output directory:
|
||||
|
||||
- Ref mode: `<REPO>/.audit-code/`
|
||||
- Local mode: `~/.claude/cache/audit-code/local-<UTC-timestamp>/`
|
||||
|
||||
Contents:
|
||||
|
||||
| File | Written by | Contents |
|
||||
|---|---|---|
|
||||
| `manifest.json` | `collect-findings.py` | full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors |
|
||||
| `manifest-<slice>.json` ×8 | `collect-findings.py` | per-agent subsets |
|
||||
| `findings-<agent>.json` ×8 | subagents | triaged findings; the walkthrough agent's file holds `overview` + `files[]` instead |
|
||||
|
||||
Ref mode additionally writes `<REPO>/audit-code-<short-ref>.md` — headline
|
||||
table, walkthrough, per-file findings by severity, dependency deltas,
|
||||
consistency observations, maintainability observations, skipped files, and
|
||||
linter coverage. Chat gets a short headline plus the path.
|
||||
|
||||
Per-tool stats in the manifest record `pre_filter` and `post_filter` counts,
|
||||
so the noise removed by diff filtering is visible per tool.
|
||||
|
||||
## Telemetry
|
||||
|
||||
Append-only JSONL at `~/.claude/cache/audit-code/runs.jsonl`. Two record
|
||||
kinds, both written by `scripts/telemetry.py`:
|
||||
|
||||
`subagent_run` — one row per agent per run: `run_id`, `repo`, `mode`, `agent`,
|
||||
`model`, `input_tokens`, `output_tokens`, `duration_ms`, `finding_count`.
|
||||
|
||||
`verdict` — one row per finding the user adjudicated: `run_id`, `agent`,
|
||||
`rule_id`, `file`, `line`, `verdict` (one of `kept`, `dismissed`,
|
||||
`false_positive`), and an optional note.
|
||||
|
||||
Runs are logged by a single call after the fan-out completes:
|
||||
|
||||
```
|
||||
echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
|
||||
| python ${SKILL_DIR}/scripts/log-run.py \
|
||||
--output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
|
||||
--mode <local|ref> --usage-json -
|
||||
```
|
||||
|
||||
`log-run.py` reads token/duration metadata from the piped JSON and reads
|
||||
`finding_count` itself by counting entries in each `findings-<agent>.json`.
|
||||
`--log-path` overrides the default log location.
|
||||
|
||||
Reading it back:
|
||||
|
||||
```
|
||||
uv run scripts/review_stats.py
|
||||
```
|
||||
|
||||
Prints a JSON blob with `runs` (distinct run count), `by_agent`, and
|
||||
`by_rule`. Per agent: cumulative tokens, duration, run count, verdict tallies,
|
||||
`precision` (kept / total adjudicated), and `tokens_per_kept`. Per
|
||||
`<agent>/<rule_id>`: verdict tallies and precision. The script takes no
|
||||
arguments and always reads the default log path.
|
||||
|
||||
Verdict capture is interactive and local-mode only. Ref mode does not
|
||||
collect verdicts, so precision figures reflect local runs alone.
|
||||
|
||||
## Development
|
||||
|
||||
```
|
||||
uv run --group dev pytest
|
||||
```
|
||||
|
||||
197 tests as of writing, covering every adapter against a recorded fixture in
|
||||
`tests/fixtures/`, plus diff filtering, git diff parsing, language detection,
|
||||
manifest serialization, slicing, package diffing, telemetry, `log-run.py`,
|
||||
`review_stats.py`, and the `collect-findings.py` CLI.
|
||||
|
||||
Tests marked `integration` require the real linters on PATH; the marker is
|
||||
declared in `pyproject.toml`.
|
||||
|
||||
Layout:
|
||||
|
||||
```
|
||||
SKILL.md orchestration procedure (source of truth)
|
||||
agents/*.md 8 subagent prompts
|
||||
scripts/collect-findings.py CLI entry point
|
||||
scripts/runner.py subprocess wrapper, tool resolution, timeouts
|
||||
scripts/git_diff.py base resolution, changed files + added-line ranges
|
||||
scripts/diff_filter.py findings → changed lines only
|
||||
scripts/language_detect.py extension → language, dep manifests, GHA paths
|
||||
scripts/manifest.py manifest dataclasses / JSON schema
|
||||
scripts/slicing.py per-agent manifest subsets
|
||||
scripts/package_diff.py requirements.txt and package.json deltas
|
||||
scripts/adapters/*.py one parser per linter → Finding[]
|
||||
scripts/rules/ bundled opengrep rules
|
||||
scripts/telemetry.py JSONL append helpers
|
||||
scripts/log-run.py post-fan-out telemetry CLI
|
||||
scripts/review_stats.py telemetry aggregation CLI
|
||||
scripts/install-tools.sh tool installer
|
||||
tests/ pytest suite + recorded tool-output fixtures
|
||||
```
|
||||
|
||||
Adding a linter means: write `scripts/adapters/<tool>.py` returning
|
||||
`Finding[]`, record a fixture of its real output, add a test, wire it into the
|
||||
right `_run_*_tools` function in `collect-findings.py`, and add its name to
|
||||
the relevant tool set in `scripts/slicing.py` so a subagent actually receives
|
||||
it.
|
||||
Reference in New Issue
Block a user