audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
and --user-only (no system package managers, no sudo).
log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.
audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).
Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
290 lines
14 KiB
Markdown
290 lines
14 KiB
Markdown
# audit-code
|
||
|
||
Automated, linter-driven audit of source-code changes. Runs an OSS SAST and
|
||
quality stack against a repo, filters every finding down to lines the change
|
||
actually touched, slices the result per reviewer concern, and fans out eight
|
||
subagents in parallel to triage.
|
||
|
||
Part of the `reviews` plugin (marketplace `mroberts`). The sibling skill
|
||
`review-pr` does the opposite job — a guided walkthrough for a human reading a
|
||
PR. `audit-code` is the tool-driven one; it produces findings, not a reading
|
||
plan.
|
||
|
||
Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell,
|
||
GitHub Actions workflows.
|
||
|
||
## When it runs
|
||
|
||
Auto-activates on phrasing like "audit this code", "run the linters on this
|
||
change", or an explicit `/audit-code`.
|
||
|
||
Two modes:
|
||
|
||
| Invocation | Mode | Diff target | Output |
|
||
|---------------------------|---------|-------------------------|---------------------------------|
|
||
| `/audit-code` | `local` | working tree vs origin | interactive walkthrough in chat |
|
||
| `/audit-code <ref>` | `ref` | `<ref>` vs origin | `audit-code-<short>.md` + chat |
|
||
| `/audit-code <pr_number>` | `ref` | PR head ref vs base | same as ref mode |
|
||
|
||
Local mode runs in the user's current checkout. Ref/PR mode isolates the
|
||
checkout first — a git worktree when the cwd is a git checkout of the target
|
||
repo, otherwise a fresh `gh repo clone` into
|
||
`~/.claude/cache/audit-code/<short-ref>/`. The clone path is what makes the
|
||
skill work from jj workspaces and from outside the repo entirely.
|
||
|
||
PR-number resolution requires `gh` on PATH and authenticated. The repo is
|
||
resolved to `owner/repo` up front and passed as `--repo` on every `gh` call,
|
||
so `gh` never autodetects from cwd.
|
||
|
||
## Pipeline
|
||
|
||
```
|
||
SKILL.md (orchestrator)
|
||
│
|
||
├─► scripts/collect-findings.py --repo --head --output-dir --mode
|
||
│ ├── resolve default branch, fetch origin, compute base
|
||
│ ├── git diff --unified=0 base...head → changed files + added-line ranges
|
||
│ ├── classify each path: supported source / GHA / dep manifest / skipped
|
||
│ ├── dispatch linters for the languages actually present
|
||
│ ├── diff-filter: drop findings not overlapping an added-line range
|
||
│ ├── write manifest.json
|
||
│ └── write 8 per-agent manifest slices
|
||
│
|
||
├─► 8 Task subagents IN PARALLEL, one per slice
|
||
│ each writes findings-<agent>.json
|
||
│
|
||
├─► aggregate: dedupe on {file, line, rule_id}, group by severity
|
||
├─► scripts/log-run.py → append telemetry rows
|
||
└─► chat summary (local) or markdown report (ref)
|
||
```
|
||
|
||
`collect-findings.py` exits non-zero when the diff cannot be computed, or when
|
||
the change set contains no supported source files and no dependency manifests.
|
||
On failure it still writes `manifest.json`; the `errors[]` array holds the
|
||
reason.
|
||
|
||
Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30
|
||
others — see `scripts/language_detect.py`) do not abort the run. Supported
|
||
files are still reviewed, and one `errors[]` entry per language records the
|
||
gap so partial coverage is visible rather than silent.
|
||
|
||
### Diff filtering
|
||
|
||
`scripts/diff_filter.py` keeps a finding only if its `[line, end_line]` span
|
||
overlaps one of the added-line ranges recorded for that file. Paths are
|
||
normalized (absolute → repo-relative, `./` stripped) before comparison, since
|
||
adapters emit paths in whichever form their tool produced. Findings in
|
||
untouched code are discarded, so the audit is scoped to the change.
|
||
|
||
### Tool resolution
|
||
|
||
`scripts/runner.py` resolves each binary against `.venv/bin/` inside the skill
|
||
directory first, then falls back to `PATH`. A tool that resolves nowhere is
|
||
recorded as `ran: false` with reason `not on PATH` and the run continues —
|
||
missing tools degrade coverage, they do not fail the audit. Every tool
|
||
invocation has a 180-second timeout.
|
||
|
||
## External tools
|
||
|
||
All of these are shelled out to. Nothing is vendored.
|
||
|
||
| Tool | Fires when | Purpose |
|
||
|---|---|---|
|
||
| `bandit` | Python in diff | Python security |
|
||
| `ruff` | Python in diff | Python lint (incl. S-rules) |
|
||
| `ruff` (idiom pass) | Python in diff | second `ruff` run, `--isolated --select SIM,PERF,UP,RET,PLR,C90,B`; reported as tool `ruff-idiom` |
|
||
| `mypy` | Python in diff | Python types |
|
||
| `eslint` | JS/TS in diff | JS/TS lint + security plugins |
|
||
| `tsc` | JS/TS in diff | TypeScript types (`--noEmit`) |
|
||
| `dotnet` | C# in diff | `dotnet build --no-incremental`; surfaces Roslyn + SecurityCodeScan (`SCS*`) |
|
||
| `selene` | Lua in diff | Lua lint |
|
||
| `luac` | Lua in diff | `luac -p` per file; syntax/parse errors |
|
||
| `pwsh` + `PSScriptAnalyzer` | PowerShell in diff | PowerShell lint; security rules split out |
|
||
| `pwsh` + `InjectionHunter` | PowerShell in diff | PowerShell injection SAST |
|
||
| `actionlint` | `.github/workflows/` or `.github/actions/` YAML in diff | workflow correctness |
|
||
| `zizmor` | same as above | workflow security (SARIF output) |
|
||
| `opengrep` | always | polyglot SAST, `--config=auto` plus bundled `scripts/rules/lua-security.yaml` |
|
||
| `gitleaks` | always | secret detection |
|
||
| `pip-audit` | always | Python dependency CVEs |
|
||
| `osv-scanner` | always | multi-ecosystem dependency CVEs |
|
||
| `lizard` | always | multi-language cyclomatic complexity |
|
||
| `vulture` | Python in diff | dead code |
|
||
| `radon` | Python in diff | complexity |
|
||
| `interrogate` | Python in diff | docstring coverage |
|
||
| `knip` (via `npx`) | JS/TS in diff | unused exports/files |
|
||
| `jscpd` | JS/TS in diff | copy-paste duplication |
|
||
|
||
Dependency manifests (`requirements*.txt`, `package.json`, lockfiles,
|
||
`*.csproj`, `*.sln`, `pyproject.toml`, …) are detected separately from source
|
||
files. `requirements*.txt` and `package.json` additionally get a structured
|
||
before/after package diff (added / removed / upgraded) computed by
|
||
`scripts/package_diff.py`.
|
||
|
||
PowerShell is worth calling out: if `pwsh` is missing, both
|
||
`psscriptanalyzer` and `injectionhunter` report unavailable. InjectionHunter
|
||
is the only injection SAST pass for PowerShell, so its absence means
|
||
PowerShell injection flaws went unchecked — that is a coverage gap, not a
|
||
clean result.
|
||
|
||
### Installing them
|
||
|
||
```
|
||
scripts/install-tools.sh # everything, including system packages
|
||
scripts/install-tools.sh --user-only # skip brew/pacman/apt (no sudo)
|
||
scripts/install-tools.sh --check-only # install nothing; non-zero if anything is missing
|
||
```
|
||
|
||
What it does:
|
||
|
||
1. `uv sync --group tools` into `${SKILL_DIR}/.venv/` — installs `bandit`,
|
||
`ruff`, `mypy`, `pip-audit`, `vulture`, `radon`, `interrogate`, `lizard`.
|
||
Falls back to `pip install bandit ruff mypy pip-audit` if `uv` is absent.
|
||
2. Downloads a platform-matched `opengrep` release binary into
|
||
`${SKILL_DIR}/.venv/bin/` (Linux x86_64/aarch64, macOS x86_64/arm64).
|
||
3. Installs the `PSScriptAnalyzer` and `InjectionHunter` PowerShell modules
|
||
for the current user, if `pwsh` is present.
|
||
4. Installs `gitleaks`, `osv-scanner`, and `gh` via Homebrew, paru/yay/pacman,
|
||
or prints apt-era install instructions.
|
||
5. Prints a resolution table for each tool, then lists the per-project tools
|
||
it deliberately does not install.
|
||
|
||
Not installed by the script — these must exist in the target repo or on PATH
|
||
yourself: `eslint` (+ `eslint-plugin-security`), `typescript`/`tsc`, `dotnet`
|
||
with `SecurityCodeScan`, `knip`, `jscpd`, `selene`, `luac`, `actionlint`,
|
||
`zizmor`.
|
||
|
||
## The eight subagents
|
||
|
||
Each subagent gets the prompt from `agents/<name>.md` plus four substituted
|
||
variables: `MANIFEST`, `REPO`, `MODE`, `OUTPUT`. Each writes
|
||
`findings-<agent>.json` into the output directory.
|
||
|
||
| Agent | Manifest slice | Covers |
|
||
|---|---|---|
|
||
| `walkthrough-reviewer` | `manifest-walkthrough.json` | Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. |
|
||
| `security-triage-reviewer` | `manifest-security-triage.json` | bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, `SCS*` from dotnet, and the PSScriptAnalyzer credential/crypto rules. |
|
||
| `type-safety-reviewer` | `manifest-type-safety.json` | mypy, tsc, non-security eslint, non-`SCS` dotnet. Also hunts leaked `Any`, unexplained `type: ignore`/`@ts-ignore`, missing hints on new public functions. |
|
||
| `dependency-reviewer` | `manifest-dependency.json` | pip-audit and osv-scanner findings, plus the `package_diffs` added/removed/upgraded entries with advisory lookups. |
|
||
| `consistency-reviewer` | `manifest-consistency.json` | Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the `ruff-idiom` findings. |
|
||
| `secrets-reviewer` | `manifest-secrets.json` | gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. |
|
||
| `maintainability-reviewer` | `manifest-maintainability.json` | vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the `C901`/`PLR0915` complexity rules from the ruff idiom pass. |
|
||
| `gha-reviewer` | `manifest-gha-reviewer.json` | actionlint and zizmor findings against changed workflow files. |
|
||
|
||
Slicing logic lives in `scripts/slicing.py`, which is the authority on which
|
||
tool's output reaches which agent.
|
||
|
||
Default models: Sonnet for seven agents, Haiku 4.5 for
|
||
`maintainability-reviewer` (its work is deterministic-tool triage). Each has
|
||
an override flag — `--security-model`, `--type-safety-model`,
|
||
`--dependency-model`, `--consistency-model`, `--secrets-model`,
|
||
`--maintainability-model`, `--walkthrough-model`, `--gha-model`.
|
||
|
||
Aggregation dedupes findings sharing `{file, line, rule_id}` — first to finish
|
||
wins, and the loser is recorded under `also_flagged_by`. Severities are
|
||
`critical`, `high`, `medium`, `low`, `info`.
|
||
|
||
## Output artifacts
|
||
|
||
Output directory:
|
||
|
||
- Ref mode: `<REPO>/.audit-code/`
|
||
- Local mode: `~/.claude/cache/audit-code/local-<UTC-timestamp>/`
|
||
|
||
Contents:
|
||
|
||
| File | Written by | Contents |
|
||
|---|---|---|
|
||
| `manifest.json` | `collect-findings.py` | full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors |
|
||
| `manifest-<slice>.json` ×8 | `collect-findings.py` | per-agent subsets |
|
||
| `findings-<agent>.json` ×8 | subagents | triaged findings; the walkthrough agent's file holds `overview` + `files[]` instead |
|
||
|
||
Ref mode additionally writes `<REPO>/audit-code-<short-ref>.md` — headline
|
||
table, walkthrough, per-file findings by severity, dependency deltas,
|
||
consistency observations, maintainability observations, skipped files, and
|
||
linter coverage. Chat gets a short headline plus the path.
|
||
|
||
Per-tool stats in the manifest record `pre_filter` and `post_filter` counts,
|
||
so the noise removed by diff filtering is visible per tool.
|
||
|
||
## Telemetry
|
||
|
||
Append-only JSONL at `~/.claude/cache/audit-code/runs.jsonl`. Two record
|
||
kinds, both written by `scripts/telemetry.py`:
|
||
|
||
`subagent_run` — one row per agent per run: `run_id`, `repo`, `mode`, `agent`,
|
||
`model`, `input_tokens`, `output_tokens`, `duration_ms`, `finding_count`.
|
||
|
||
`verdict` — one row per finding the user adjudicated: `run_id`, `agent`,
|
||
`rule_id`, `file`, `line`, `verdict` (one of `kept`, `dismissed`,
|
||
`false_positive`), and an optional note.
|
||
|
||
Runs are logged by a single call after the fan-out completes:
|
||
|
||
```
|
||
echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
|
||
| uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
|
||
--output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
|
||
--mode <local|ref> --usage-json -
|
||
```
|
||
|
||
`log-run.py` reads token/duration metadata from the piped JSON and reads
|
||
`finding_count` itself by counting entries in each `findings-<agent>.json`.
|
||
`--log-path` overrides the default log location.
|
||
|
||
Reading it back:
|
||
|
||
```
|
||
uv run scripts/review_stats.py
|
||
```
|
||
|
||
Prints a JSON blob with `runs` (distinct run count), `by_agent`, and
|
||
`by_rule`. Per agent: cumulative tokens, duration, run count, verdict tallies,
|
||
`precision` (kept / total adjudicated), and `tokens_per_kept`. Per
|
||
`<agent>/<rule_id>`: verdict tallies and precision. The script takes no
|
||
arguments and always reads the default log path.
|
||
|
||
Verdict capture is interactive and local-mode only. Ref mode does not
|
||
collect verdicts, so precision figures reflect local runs alone.
|
||
|
||
## Development
|
||
|
||
```
|
||
uv run --group dev pytest
|
||
```
|
||
|
||
197 tests as of writing, covering every adapter against a recorded fixture in
|
||
`tests/fixtures/`, plus diff filtering, git diff parsing, language detection,
|
||
manifest serialization, slicing, package diffing, telemetry, `log-run.py`,
|
||
`review_stats.py`, and the `collect-findings.py` CLI.
|
||
|
||
Tests marked `integration` require the real linters on PATH; the marker is
|
||
declared in `pyproject.toml`.
|
||
|
||
Layout:
|
||
|
||
```
|
||
SKILL.md orchestration procedure (source of truth)
|
||
agents/*.md 8 subagent prompts
|
||
scripts/collect-findings.py CLI entry point
|
||
scripts/runner.py subprocess wrapper, tool resolution, timeouts
|
||
scripts/git_diff.py base resolution, changed files + added-line ranges
|
||
scripts/diff_filter.py findings → changed lines only
|
||
scripts/language_detect.py extension → language, dep manifests, GHA paths
|
||
scripts/manifest.py manifest dataclasses / JSON schema
|
||
scripts/slicing.py per-agent manifest subsets
|
||
scripts/package_diff.py requirements.txt and package.json deltas
|
||
scripts/adapters/*.py one parser per linter → Finding[]
|
||
scripts/rules/ bundled opengrep rules
|
||
scripts/telemetry.py JSONL append helpers
|
||
scripts/log-run.py post-fan-out telemetry CLI
|
||
scripts/review_stats.py telemetry aggregation CLI
|
||
scripts/install-tools.sh tool installer
|
||
tests/ pytest suite + recorded tool-output fixtures
|
||
```
|
||
|
||
Adding a linter means: write `scripts/adapters/<tool>.py` returning
|
||
`Finding[]`, record a fixture of its real output, add a test, wire it into the
|
||
right `_run_*_tools` function in `collect-findings.py`, and add its name to
|
||
the relevant tool set in `scripts/slicing.py` so a subagent actually receives
|
||
it.
|