Files
claude-plugin/plugins/reviews/skills/audit-code/README.md
T
mroberts 37fc3fb291 Fix audit tool bootstrap and add per-run preflight
audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
  under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
  since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
  and --user-only (no system package managers, no sudo).

log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.

audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).

Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
2026-09-22 15:21:28 -05:00

290 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# audit-code
Automated, linter-driven audit of source-code changes. Runs an OSS SAST and
quality stack against a repo, filters every finding down to lines the change
actually touched, slices the result per reviewer concern, and fans out eight
subagents in parallel to triage.
Part of the `reviews` plugin (marketplace `mroberts`). The sibling skill
`review-pr` does the opposite job — a guided walkthrough for a human reading a
PR. `audit-code` is the tool-driven one; it produces findings, not a reading
plan.
Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell,
GitHub Actions workflows.
## When it runs
Auto-activates on phrasing like "audit this code", "run the linters on this
change", or an explicit `/audit-code`.
Two modes:
| Invocation | Mode | Diff target | Output |
|---------------------------|---------|-------------------------|---------------------------------|
| `/audit-code` | `local` | working tree vs origin | interactive walkthrough in chat |
| `/audit-code <ref>` | `ref` | `<ref>` vs origin | `audit-code-<short>.md` + chat |
| `/audit-code <pr_number>` | `ref` | PR head ref vs base | same as ref mode |
Local mode runs in the user's current checkout. Ref/PR mode isolates the
checkout first — a git worktree when the cwd is a git checkout of the target
repo, otherwise a fresh `gh repo clone` into
`~/.claude/cache/audit-code/<short-ref>/`. The clone path is what makes the
skill work from jj workspaces and from outside the repo entirely.
PR-number resolution requires `gh` on PATH and authenticated. The repo is
resolved to `owner/repo` up front and passed as `--repo` on every `gh` call,
so `gh` never autodetects from cwd.
## Pipeline
```
SKILL.md (orchestrator)
│
├─► scripts/collect-findings.py --repo --head --output-dir --mode
│ ├── resolve default branch, fetch origin, compute base
│ ├── git diff --unified=0 base...head → changed files + added-line ranges
│ ├── classify each path: supported source / GHA / dep manifest / skipped
│ ├── dispatch linters for the languages actually present
│ ├── diff-filter: drop findings not overlapping an added-line range
│ ├── write manifest.json
│ └── write 8 per-agent manifest slices
│
├─► 8 Task subagents IN PARALLEL, one per slice
│ each writes findings-<agent>.json
│
├─► aggregate: dedupe on {file, line, rule_id}, group by severity
├─► scripts/log-run.py → append telemetry rows
└─► chat summary (local) or markdown report (ref)
```
`collect-findings.py` exits non-zero when the diff cannot be computed, or when
the change set contains no supported source files and no dependency manifests.
On failure it still writes `manifest.json`; the `errors[]` array holds the
reason.
Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30
others — see `scripts/language_detect.py`) do not abort the run. Supported
files are still reviewed, and one `errors[]` entry per language records the
gap so partial coverage is visible rather than silent.
### Diff filtering
`scripts/diff_filter.py` keeps a finding only if its `[line, end_line]` span
overlaps one of the added-line ranges recorded for that file. Paths are
normalized (absolute → repo-relative, `./` stripped) before comparison, since
adapters emit paths in whichever form their tool produced. Findings in
untouched code are discarded, so the audit is scoped to the change.
### Tool resolution
`scripts/runner.py` resolves each binary against `.venv/bin/` inside the skill
directory first, then falls back to `PATH`. A tool that resolves nowhere is
recorded as `ran: false` with reason `not on PATH` and the run continues —
missing tools degrade coverage, they do not fail the audit. Every tool
invocation has a 180-second timeout.
## External tools
All of these are shelled out to. Nothing is vendored.
| Tool | Fires when | Purpose |
|---|---|---|
| `bandit` | Python in diff | Python security |
| `ruff` | Python in diff | Python lint (incl. S-rules) |
| `ruff` (idiom pass) | Python in diff | second `ruff` run, `--isolated --select SIM,PERF,UP,RET,PLR,C90,B`; reported as tool `ruff-idiom` |
| `mypy` | Python in diff | Python types |
| `eslint` | JS/TS in diff | JS/TS lint + security plugins |
| `tsc` | JS/TS in diff | TypeScript types (`--noEmit`) |
| `dotnet` | C# in diff | `dotnet build --no-incremental`; surfaces Roslyn + SecurityCodeScan (`SCS*`) |
| `selene` | Lua in diff | Lua lint |
| `luac` | Lua in diff | `luac -p` per file; syntax/parse errors |
| `pwsh` + `PSScriptAnalyzer` | PowerShell in diff | PowerShell lint; security rules split out |
| `pwsh` + `InjectionHunter` | PowerShell in diff | PowerShell injection SAST |
| `actionlint` | `.github/workflows/` or `.github/actions/` YAML in diff | workflow correctness |
| `zizmor` | same as above | workflow security (SARIF output) |
| `opengrep` | always | polyglot SAST, `--config=auto` plus bundled `scripts/rules/lua-security.yaml` |
| `gitleaks` | always | secret detection |
| `pip-audit` | always | Python dependency CVEs |
| `osv-scanner` | always | multi-ecosystem dependency CVEs |
| `lizard` | always | multi-language cyclomatic complexity |
| `vulture` | Python in diff | dead code |
| `radon` | Python in diff | complexity |
| `interrogate` | Python in diff | docstring coverage |
| `knip` (via `npx`) | JS/TS in diff | unused exports/files |
| `jscpd` | JS/TS in diff | copy-paste duplication |
Dependency manifests (`requirements*.txt`, `package.json`, lockfiles,
`*.csproj`, `*.sln`, `pyproject.toml`, …) are detected separately from source
files. `requirements*.txt` and `package.json` additionally get a structured
before/after package diff (added / removed / upgraded) computed by
`scripts/package_diff.py`.
PowerShell is worth calling out: if `pwsh` is missing, both
`psscriptanalyzer` and `injectionhunter` report unavailable. InjectionHunter
is the only injection SAST pass for PowerShell, so its absence means
PowerShell injection flaws went unchecked — that is a coverage gap, not a
clean result.
### Installing them
```
scripts/install-tools.sh # everything, including system packages
scripts/install-tools.sh --user-only # skip brew/pacman/apt (no sudo)
scripts/install-tools.sh --check-only # install nothing; non-zero if anything is missing
```
What it does:
1. `uv sync --group tools` into `${SKILL_DIR}/.venv/` — installs `bandit`,
`ruff`, `mypy`, `pip-audit`, `vulture`, `radon`, `interrogate`, `lizard`.
Falls back to `pip install bandit ruff mypy pip-audit` if `uv` is absent.
2. Downloads a platform-matched `opengrep` release binary into
`${SKILL_DIR}/.venv/bin/` (Linux x86_64/aarch64, macOS x86_64/arm64).
3. Installs the `PSScriptAnalyzer` and `InjectionHunter` PowerShell modules
for the current user, if `pwsh` is present.
4. Installs `gitleaks`, `osv-scanner`, and `gh` via Homebrew, paru/yay/pacman,
or prints apt-era install instructions.
5. Prints a resolution table for each tool, then lists the per-project tools
it deliberately does not install.
Not installed by the script — these must exist in the target repo or on PATH
yourself: `eslint` (+ `eslint-plugin-security`), `typescript`/`tsc`, `dotnet`
with `SecurityCodeScan`, `knip`, `jscpd`, `selene`, `luac`, `actionlint`,
`zizmor`.
## The eight subagents
Each subagent gets the prompt from `agents/<name>.md` plus four substituted
variables: `MANIFEST`, `REPO`, `MODE`, `OUTPUT`. Each writes
`findings-<agent>.json` into the output directory.
| Agent | Manifest slice | Covers |
|---|---|---|
| `walkthrough-reviewer` | `manifest-walkthrough.json` | Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. |
| `security-triage-reviewer` | `manifest-security-triage.json` | bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, `SCS*` from dotnet, and the PSScriptAnalyzer credential/crypto rules. |
| `type-safety-reviewer` | `manifest-type-safety.json` | mypy, tsc, non-security eslint, non-`SCS` dotnet. Also hunts leaked `Any`, unexplained `type: ignore`/`@ts-ignore`, missing hints on new public functions. |
| `dependency-reviewer` | `manifest-dependency.json` | pip-audit and osv-scanner findings, plus the `package_diffs` added/removed/upgraded entries with advisory lookups. |
| `consistency-reviewer` | `manifest-consistency.json` | Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the `ruff-idiom` findings. |
| `secrets-reviewer` | `manifest-secrets.json` | gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. |
| `maintainability-reviewer` | `manifest-maintainability.json` | vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the `C901`/`PLR0915` complexity rules from the ruff idiom pass. |
| `gha-reviewer` | `manifest-gha-reviewer.json` | actionlint and zizmor findings against changed workflow files. |
Slicing logic lives in `scripts/slicing.py`, which is the authority on which
tool's output reaches which agent.
Default models: Sonnet for seven agents, Haiku 4.5 for
`maintainability-reviewer` (its work is deterministic-tool triage). Each has
an override flag — `--security-model`, `--type-safety-model`,
`--dependency-model`, `--consistency-model`, `--secrets-model`,
`--maintainability-model`, `--walkthrough-model`, `--gha-model`.
Aggregation dedupes findings sharing `{file, line, rule_id}` — first to finish
wins, and the loser is recorded under `also_flagged_by`. Severities are
`critical`, `high`, `medium`, `low`, `info`.
## Output artifacts
Output directory:
- Ref mode: `<REPO>/.audit-code/`
- Local mode: `~/.claude/cache/audit-code/local-<UTC-timestamp>/`
Contents:
| File | Written by | Contents |
|---|---|---|
| `manifest.json` | `collect-findings.py` | full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors |
| `manifest-<slice>.json` ×8 | `collect-findings.py` | per-agent subsets |
| `findings-<agent>.json` ×8 | subagents | triaged findings; the walkthrough agent's file holds `overview` + `files[]` instead |
Ref mode additionally writes `<REPO>/audit-code-<short-ref>.md` — headline
table, walkthrough, per-file findings by severity, dependency deltas,
consistency observations, maintainability observations, skipped files, and
linter coverage. Chat gets a short headline plus the path.
Per-tool stats in the manifest record `pre_filter` and `post_filter` counts,
so the noise removed by diff filtering is visible per tool.
## Telemetry
Append-only JSONL at `~/.claude/cache/audit-code/runs.jsonl`. Two record
kinds, both written by `scripts/telemetry.py`:
`subagent_run` — one row per agent per run: `run_id`, `repo`, `mode`, `agent`,
`model`, `input_tokens`, `output_tokens`, `duration_ms`, `finding_count`.
`verdict` — one row per finding the user adjudicated: `run_id`, `agent`,
`rule_id`, `file`, `line`, `verdict` (one of `kept`, `dismissed`,
`false_positive`), and an optional note.
Runs are logged by a single call after the fan-out completes:
```
echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
| uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
--mode <local|ref> --usage-json -
```
`log-run.py` reads token/duration metadata from the piped JSON and reads
`finding_count` itself by counting entries in each `findings-<agent>.json`.
`--log-path` overrides the default log location.
Reading it back:
```
uv run scripts/review_stats.py
```
Prints a JSON blob with `runs` (distinct run count), `by_agent`, and
`by_rule`. Per agent: cumulative tokens, duration, run count, verdict tallies,
`precision` (kept / total adjudicated), and `tokens_per_kept`. Per
`<agent>/<rule_id>`: verdict tallies and precision. The script takes no
arguments and always reads the default log path.
Verdict capture is interactive and local-mode only. Ref mode does not
collect verdicts, so precision figures reflect local runs alone.
## Development
```
uv run --group dev pytest
```
197 tests as of writing, covering every adapter against a recorded fixture in
`tests/fixtures/`, plus diff filtering, git diff parsing, language detection,
manifest serialization, slicing, package diffing, telemetry, `log-run.py`,
`review_stats.py`, and the `collect-findings.py` CLI.
Tests marked `integration` require the real linters on PATH; the marker is
declared in `pyproject.toml`.
Layout:
```
SKILL.md orchestration procedure (source of truth)
agents/*.md 8 subagent prompts
scripts/collect-findings.py CLI entry point
scripts/runner.py subprocess wrapper, tool resolution, timeouts
scripts/git_diff.py base resolution, changed files + added-line ranges
scripts/diff_filter.py findings → changed lines only
scripts/language_detect.py extension → language, dep manifests, GHA paths
scripts/manifest.py manifest dataclasses / JSON schema
scripts/slicing.py per-agent manifest subsets
scripts/package_diff.py requirements.txt and package.json deltas
scripts/adapters/*.py one parser per linter → Finding[]
scripts/rules/ bundled opengrep rules
scripts/telemetry.py JSONL append helpers
scripts/log-run.py post-fan-out telemetry CLI
scripts/review_stats.py telemetry aggregation CLI
scripts/install-tools.sh tool installer
tests/ pytest suite + recorded tool-output fixtures
```
Adding a linter means: write `scripts/adapters/<tool>.py` returning
`Finding[]`, record a fixture of its real output, add a test, wire it into the
right `_run_*_tools` function in `collect-findings.py`, and add its name to
the relevant tool set in `scripts/slicing.py` so a subagent actually receives
it.