audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
and --user-only (no system package managers, no sudo).
log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.
audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).
Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
14 KiB
audit-code
Automated, linter-driven audit of source-code changes. Runs an OSS SAST and quality stack against a repo, filters every finding down to lines the change actually touched, slices the result per reviewer concern, and fans out eight subagents in parallel to triage.
Part of the reviews plugin (marketplace mroberts). The sibling skill
review-pr does the opposite job — a guided walkthrough for a human reading a
PR. audit-code is the tool-driven one; it produces findings, not a reading
plan.
Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions workflows.
When it runs
Auto-activates on phrasing like "audit this code", "run the linters on this
change", or an explicit /audit-code.
Two modes:
| Invocation | Mode | Diff target | Output |
|---|---|---|---|
/audit-code |
local |
working tree vs origin | interactive walkthrough in chat |
/audit-code <ref> |
ref |
<ref> vs origin |
audit-code-<short>.md + chat |
/audit-code <pr_number> |
ref |
PR head ref vs base | same as ref mode |
Local mode runs in the user's current checkout. Ref/PR mode isolates the
checkout first — a git worktree when the cwd is a git checkout of the target
repo, otherwise a fresh gh repo clone into
~/.claude/cache/audit-code/<short-ref>/. The clone path is what makes the
skill work from jj workspaces and from outside the repo entirely.
PR-number resolution requires gh on PATH and authenticated. The repo is
resolved to owner/repo up front and passed as --repo on every gh call,
so gh never autodetects from cwd.
Pipeline
SKILL.md (orchestrator)
│
├─► scripts/collect-findings.py --repo --head --output-dir --mode
│ ├── resolve default branch, fetch origin, compute base
│ ├── git diff --unified=0 base...head → changed files + added-line ranges
│ ├── classify each path: supported source / GHA / dep manifest / skipped
│ ├── dispatch linters for the languages actually present
│ ├── diff-filter: drop findings not overlapping an added-line range
│ ├── write manifest.json
│ └── write 8 per-agent manifest slices
│
├─► 8 Task subagents IN PARALLEL, one per slice
│ each writes findings-<agent>.json
│
├─► aggregate: dedupe on {file, line, rule_id}, group by severity
├─► scripts/log-run.py → append telemetry rows
└─► chat summary (local) or markdown report (ref)
collect-findings.py exits non-zero when the diff cannot be computed, or when
the change set contains no supported source files and no dependency manifests.
On failure it still writes manifest.json; the errors[] array holds the
reason.
Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30
others — see scripts/language_detect.py) do not abort the run. Supported
files are still reviewed, and one errors[] entry per language records the
gap so partial coverage is visible rather than silent.
Diff filtering
scripts/diff_filter.py keeps a finding only if its [line, end_line] span
overlaps one of the added-line ranges recorded for that file. Paths are
normalized (absolute → repo-relative, ./ stripped) before comparison, since
adapters emit paths in whichever form their tool produced. Findings in
untouched code are discarded, so the audit is scoped to the change.
Tool resolution
scripts/runner.py resolves each binary against .venv/bin/ inside the skill
directory first, then falls back to PATH. A tool that resolves nowhere is
recorded as ran: false with reason not on PATH and the run continues —
missing tools degrade coverage, they do not fail the audit. Every tool
invocation has a 180-second timeout.
External tools
All of these are shelled out to. Nothing is vendored.
| Tool | Fires when | Purpose |
|---|---|---|
bandit |
Python in diff | Python security |
ruff |
Python in diff | Python lint (incl. S-rules) |
ruff (idiom pass) |
Python in diff | second ruff run, --isolated --select SIM,PERF,UP,RET,PLR,C90,B; reported as tool ruff-idiom |
mypy |
Python in diff | Python types |
eslint |
JS/TS in diff | JS/TS lint + security plugins |
tsc |
JS/TS in diff | TypeScript types (--noEmit) |
dotnet |
C# in diff | dotnet build --no-incremental; surfaces Roslyn + SecurityCodeScan (SCS*) |
selene |
Lua in diff | Lua lint |
luac |
Lua in diff | luac -p per file; syntax/parse errors |
pwsh + PSScriptAnalyzer |
PowerShell in diff | PowerShell lint; security rules split out |
pwsh + InjectionHunter |
PowerShell in diff | PowerShell injection SAST |
actionlint |
.github/workflows/ or .github/actions/ YAML in diff |
workflow correctness |
zizmor |
same as above | workflow security (SARIF output) |
opengrep |
always | polyglot SAST, --config=auto plus bundled scripts/rules/lua-security.yaml |
gitleaks |
always | secret detection |
pip-audit |
always | Python dependency CVEs |
osv-scanner |
always | multi-ecosystem dependency CVEs |
lizard |
always | multi-language cyclomatic complexity |
vulture |
Python in diff | dead code |
radon |
Python in diff | complexity |
interrogate |
Python in diff | docstring coverage |
knip (via npx) |
JS/TS in diff | unused exports/files |
jscpd |
JS/TS in diff | copy-paste duplication |
Dependency manifests (requirements*.txt, package.json, lockfiles,
*.csproj, *.sln, pyproject.toml, …) are detected separately from source
files. requirements*.txt and package.json additionally get a structured
before/after package diff (added / removed / upgraded) computed by
scripts/package_diff.py.
PowerShell is worth calling out: if pwsh is missing, both
psscriptanalyzer and injectionhunter report unavailable. InjectionHunter
is the only injection SAST pass for PowerShell, so its absence means
PowerShell injection flaws went unchecked — that is a coverage gap, not a
clean result.
Installing them
scripts/install-tools.sh # everything, including system packages
scripts/install-tools.sh --user-only # skip brew/pacman/apt (no sudo)
scripts/install-tools.sh --check-only # install nothing; non-zero if anything is missing
What it does:
uv sync --group toolsinto${SKILL_DIR}/.venv/— installsbandit,ruff,mypy,pip-audit,vulture,radon,interrogate,lizard. Falls back topip install bandit ruff mypy pip-auditifuvis absent.- Downloads a platform-matched
opengreprelease binary into${SKILL_DIR}/.venv/bin/(Linux x86_64/aarch64, macOS x86_64/arm64). - Installs the
PSScriptAnalyzerandInjectionHunterPowerShell modules for the current user, ifpwshis present. - Installs
gitleaks,osv-scanner, andghvia Homebrew, paru/yay/pacman, or prints apt-era install instructions. - Prints a resolution table for each tool, then lists the per-project tools it deliberately does not install.
Not installed by the script — these must exist in the target repo or on PATH
yourself: eslint (+ eslint-plugin-security), typescript/tsc, dotnet
with SecurityCodeScan, knip, jscpd, selene, luac, actionlint,
zizmor.
The eight subagents
Each subagent gets the prompt from agents/<name>.md plus four substituted
variables: MANIFEST, REPO, MODE, OUTPUT. Each writes
findings-<agent>.json into the output directory.
| Agent | Manifest slice | Covers |
|---|---|---|
walkthrough-reviewer |
manifest-walkthrough.json |
Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. |
security-triage-reviewer |
manifest-security-triage.json |
bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, SCS* from dotnet, and the PSScriptAnalyzer credential/crypto rules. |
type-safety-reviewer |
manifest-type-safety.json |
mypy, tsc, non-security eslint, non-SCS dotnet. Also hunts leaked Any, unexplained type: ignore/@ts-ignore, missing hints on new public functions. |
dependency-reviewer |
manifest-dependency.json |
pip-audit and osv-scanner findings, plus the package_diffs added/removed/upgraded entries with advisory lookups. |
consistency-reviewer |
manifest-consistency.json |
Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the ruff-idiom findings. |
secrets-reviewer |
manifest-secrets.json |
gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. |
maintainability-reviewer |
manifest-maintainability.json |
vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the C901/PLR0915 complexity rules from the ruff idiom pass. |
gha-reviewer |
manifest-gha-reviewer.json |
actionlint and zizmor findings against changed workflow files. |
Slicing logic lives in scripts/slicing.py, which is the authority on which
tool's output reaches which agent.
Default models: Sonnet for seven agents, Haiku 4.5 for
maintainability-reviewer (its work is deterministic-tool triage). Each has
an override flag — --security-model, --type-safety-model,
--dependency-model, --consistency-model, --secrets-model,
--maintainability-model, --walkthrough-model, --gha-model.
Aggregation dedupes findings sharing {file, line, rule_id} — first to finish
wins, and the loser is recorded under also_flagged_by. Severities are
critical, high, medium, low, info.
Output artifacts
Output directory:
- Ref mode:
<REPO>/.audit-code/ - Local mode:
~/.claude/cache/audit-code/local-<UTC-timestamp>/
Contents:
| File | Written by | Contents |
|---|---|---|
manifest.json |
collect-findings.py |
full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors |
manifest-<slice>.json ×8 |
collect-findings.py |
per-agent subsets |
findings-<agent>.json ×8 |
subagents | triaged findings; the walkthrough agent's file holds overview + files[] instead |
Ref mode additionally writes <REPO>/audit-code-<short-ref>.md — headline
table, walkthrough, per-file findings by severity, dependency deltas,
consistency observations, maintainability observations, skipped files, and
linter coverage. Chat gets a short headline plus the path.
Per-tool stats in the manifest record pre_filter and post_filter counts,
so the noise removed by diff filtering is visible per tool.
Telemetry
Append-only JSONL at ~/.claude/cache/audit-code/runs.jsonl. Two record
kinds, both written by scripts/telemetry.py:
subagent_run — one row per agent per run: run_id, repo, mode, agent,
model, input_tokens, output_tokens, duration_ms, finding_count.
verdict — one row per finding the user adjudicated: run_id, agent,
rule_id, file, line, verdict (one of kept, dismissed,
false_positive), and an optional note.
Runs are logged by a single call after the fan-out completes:
echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
| uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
--mode <local|ref> --usage-json -
log-run.py reads token/duration metadata from the piped JSON and reads
finding_count itself by counting entries in each findings-<agent>.json.
--log-path overrides the default log location.
Reading it back:
uv run scripts/review_stats.py
Prints a JSON blob with runs (distinct run count), by_agent, and
by_rule. Per agent: cumulative tokens, duration, run count, verdict tallies,
precision (kept / total adjudicated), and tokens_per_kept. Per
<agent>/<rule_id>: verdict tallies and precision. The script takes no
arguments and always reads the default log path.
Verdict capture is interactive and local-mode only. Ref mode does not collect verdicts, so precision figures reflect local runs alone.
Development
uv run --group dev pytest
197 tests as of writing, covering every adapter against a recorded fixture in
tests/fixtures/, plus diff filtering, git diff parsing, language detection,
manifest serialization, slicing, package diffing, telemetry, log-run.py,
review_stats.py, and the collect-findings.py CLI.
Tests marked integration require the real linters on PATH; the marker is
declared in pyproject.toml.
Layout:
SKILL.md orchestration procedure (source of truth)
agents/*.md 8 subagent prompts
scripts/collect-findings.py CLI entry point
scripts/runner.py subprocess wrapper, tool resolution, timeouts
scripts/git_diff.py base resolution, changed files + added-line ranges
scripts/diff_filter.py findings → changed lines only
scripts/language_detect.py extension → language, dep manifests, GHA paths
scripts/manifest.py manifest dataclasses / JSON schema
scripts/slicing.py per-agent manifest subsets
scripts/package_diff.py requirements.txt and package.json deltas
scripts/adapters/*.py one parser per linter → Finding[]
scripts/rules/ bundled opengrep rules
scripts/telemetry.py JSONL append helpers
scripts/log-run.py post-fan-out telemetry CLI
scripts/review_stats.py telemetry aggregation CLI
scripts/install-tools.sh tool installer
tests/ pytest suite + recorded tool-output fixtures
Adding a linter means: write scripts/adapters/<tool>.py returning
Finding[], record a fixture of its real output, add a test, wire it into the
right _run_*_tools function in collect-findings.py, and add its name to
the relevant tool set in scripts/slicing.py so a subagent actually receives
it.