# audit-code Automated, linter-driven audit of source-code changes. Runs an OSS SAST and quality stack against a repo, filters every finding down to lines the change actually touched, slices the result per reviewer concern, and fans out eight subagents in parallel to triage. Part of the `reviews` plugin (marketplace `mroberts`). The sibling skill `review-pr` does the opposite job — a guided walkthrough for a human reading a PR. `audit-code` is the tool-driven one; it produces findings, not a reading plan. Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions workflows. ## When it runs Auto-activates on phrasing like "audit this code", "run the linters on this change", or an explicit `/audit-code`. Two modes: | Invocation | Mode | Diff target | Output | |---------------------------|---------|-------------------------|---------------------------------| | `/audit-code` | `local` | working tree vs origin | interactive walkthrough in chat | | `/audit-code ` | `ref` | `` vs origin | `audit-code-.md` + chat | | `/audit-code ` | `ref` | PR head ref vs base | same as ref mode | Local mode runs in the user's current checkout. Ref/PR mode isolates the checkout first — a git worktree when the cwd is a git checkout of the target repo, otherwise a fresh `gh repo clone` into `~/.claude/cache/audit-code//`. The clone path is what makes the skill work from jj workspaces and from outside the repo entirely. PR-number resolution requires `gh` on PATH and authenticated. The repo is resolved to `owner/repo` up front and passed as `--repo` on every `gh` call, so `gh` never autodetects from cwd. ## Pipeline ``` SKILL.md (orchestrator) │ ├─► scripts/collect-findings.py --repo --head --output-dir --mode │ ├── resolve default branch, fetch origin, compute base │ ├── git diff --unified=0 base...head → changed files + added-line ranges │ ├── classify each path: supported source / GHA / dep manifest / skipped │ ├── dispatch linters for the languages actually present │ ├── diff-filter: drop findings not overlapping an added-line range │ ├── write manifest.json │ └── write 8 per-agent manifest slices │ ├─► 8 Task subagents IN PARALLEL, one per slice │ each writes findings-.json │ ├─► aggregate: dedupe on {file, line, rule_id}, group by severity ├─► scripts/log-run.py → append telemetry rows └─► chat summary (local) or markdown report (ref) ``` `collect-findings.py` exits non-zero when the diff cannot be computed, or when the change set contains no supported source files and no dependency manifests. On failure it still writes `manifest.json`; the `errors[]` array holds the reason. Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30 others — see `scripts/language_detect.py`) do not abort the run. Supported files are still reviewed, and one `errors[]` entry per language records the gap so partial coverage is visible rather than silent. ### Diff filtering `scripts/diff_filter.py` keeps a finding only if its `[line, end_line]` span overlaps one of the added-line ranges recorded for that file. Paths are normalized (absolute → repo-relative, `./` stripped) before comparison, since adapters emit paths in whichever form their tool produced. Findings in untouched code are discarded, so the audit is scoped to the change. ### Tool resolution `scripts/runner.py` resolves each binary against `.venv/bin/` inside the skill directory first, then falls back to `PATH`. A tool that resolves nowhere is recorded as `ran: false` with reason `not on PATH` and the run continues — missing tools degrade coverage, they do not fail the audit. Every tool invocation has a 180-second timeout. ## External tools All of these are shelled out to. Nothing is vendored. | Tool | Fires when | Purpose | |---|---|---| | `bandit` | Python in diff | Python security | | `ruff` | Python in diff | Python lint (incl. S-rules) | | `ruff` (idiom pass) | Python in diff | second `ruff` run, `--isolated --select SIM,PERF,UP,RET,PLR,C90,B`; reported as tool `ruff-idiom` | | `mypy` | Python in diff | Python types | | `eslint` | JS/TS in diff | JS/TS lint + security plugins | | `tsc` | JS/TS in diff | TypeScript types (`--noEmit`) | | `dotnet` | C# in diff | `dotnet build --no-incremental`; surfaces Roslyn + SecurityCodeScan (`SCS*`) | | `selene` | Lua in diff | Lua lint | | `luac` | Lua in diff | `luac -p` per file; syntax/parse errors | | `pwsh` + `PSScriptAnalyzer` | PowerShell in diff | PowerShell lint; security rules split out | | `pwsh` + `InjectionHunter` | PowerShell in diff | PowerShell injection SAST | | `actionlint` | `.github/workflows/` or `.github/actions/` YAML in diff | workflow correctness | | `zizmor` | same as above | workflow security (SARIF output) | | `opengrep` | always | polyglot SAST, `--config=auto` plus bundled `scripts/rules/lua-security.yaml` | | `gitleaks` | always | secret detection | | `pip-audit` | always | Python dependency CVEs | | `osv-scanner` | always | multi-ecosystem dependency CVEs | | `lizard` | always | multi-language cyclomatic complexity | | `vulture` | Python in diff | dead code | | `radon` | Python in diff | complexity | | `interrogate` | Python in diff | docstring coverage | | `knip` (via `npx`) | JS/TS in diff | unused exports/files | | `jscpd` | JS/TS in diff | copy-paste duplication | Dependency manifests (`requirements*.txt`, `package.json`, lockfiles, `*.csproj`, `*.sln`, `pyproject.toml`, …) are detected separately from source files. `requirements*.txt` and `package.json` additionally get a structured before/after package diff (added / removed / upgraded) computed by `scripts/package_diff.py`. PowerShell is worth calling out: if `pwsh` is missing, both `psscriptanalyzer` and `injectionhunter` report unavailable. InjectionHunter is the only injection SAST pass for PowerShell, so its absence means PowerShell injection flaws went unchecked — that is a coverage gap, not a clean result. ### Installing them ``` scripts/install-tools.sh # everything, including system packages scripts/install-tools.sh --user-only # skip brew/pacman/apt (no sudo) scripts/install-tools.sh --check-only # install nothing; non-zero if anything is missing ``` What it does: 1. `uv sync --group tools` into `${SKILL_DIR}/.venv/` — installs `bandit`, `ruff`, `mypy`, `pip-audit`, `vulture`, `radon`, `interrogate`, `lizard`. Falls back to `pip install bandit ruff mypy pip-audit` if `uv` is absent. 2. Downloads a platform-matched `opengrep` release binary into `${SKILL_DIR}/.venv/bin/` (Linux x86_64/aarch64, macOS x86_64/arm64). 3. Installs the `PSScriptAnalyzer` and `InjectionHunter` PowerShell modules for the current user, if `pwsh` is present. 4. Installs `gitleaks`, `osv-scanner`, and `gh` via Homebrew, paru/yay/pacman, or prints apt-era install instructions. 5. Prints a resolution table for each tool, then lists the per-project tools it deliberately does not install. Not installed by the script — these must exist in the target repo or on PATH yourself: `eslint` (+ `eslint-plugin-security`), `typescript`/`tsc`, `dotnet` with `SecurityCodeScan`, `knip`, `jscpd`, `selene`, `luac`, `actionlint`, `zizmor`. ## The eight subagents Each subagent gets the prompt from `agents/.md` plus four substituted variables: `MANIFEST`, `REPO`, `MODE`, `OUTPUT`. Each writes `findings-.json` into the output directory. | Agent | Manifest slice | Covers | |---|---|---| | `walkthrough-reviewer` | `manifest-walkthrough.json` | Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. | | `security-triage-reviewer` | `manifest-security-triage.json` | bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, `SCS*` from dotnet, and the PSScriptAnalyzer credential/crypto rules. | | `type-safety-reviewer` | `manifest-type-safety.json` | mypy, tsc, non-security eslint, non-`SCS` dotnet. Also hunts leaked `Any`, unexplained `type: ignore`/`@ts-ignore`, missing hints on new public functions. | | `dependency-reviewer` | `manifest-dependency.json` | pip-audit and osv-scanner findings, plus the `package_diffs` added/removed/upgraded entries with advisory lookups. | | `consistency-reviewer` | `manifest-consistency.json` | Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the `ruff-idiom` findings. | | `secrets-reviewer` | `manifest-secrets.json` | gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. | | `maintainability-reviewer` | `manifest-maintainability.json` | vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the `C901`/`PLR0915` complexity rules from the ruff idiom pass. | | `gha-reviewer` | `manifest-gha-reviewer.json` | actionlint and zizmor findings against changed workflow files. | Slicing logic lives in `scripts/slicing.py`, which is the authority on which tool's output reaches which agent. Default models: Sonnet for seven agents, Haiku 4.5 for `maintainability-reviewer` (its work is deterministic-tool triage). Each has an override flag — `--security-model`, `--type-safety-model`, `--dependency-model`, `--consistency-model`, `--secrets-model`, `--maintainability-model`, `--walkthrough-model`, `--gha-model`. Aggregation dedupes findings sharing `{file, line, rule_id}` — first to finish wins, and the loser is recorded under `also_flagged_by`. Severities are `critical`, `high`, `medium`, `low`, `info`. ## Output artifacts Output directory: - Ref mode: `/.audit-code/` - Local mode: `~/.claude/cache/audit-code/local-/` Contents: | File | Written by | Contents | |---|---|---| | `manifest.json` | `collect-findings.py` | full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors | | `manifest-.json` ×8 | `collect-findings.py` | per-agent subsets | | `findings-.json` ×8 | subagents | triaged findings; the walkthrough agent's file holds `overview` + `files[]` instead | Ref mode additionally writes `/audit-code-.md` — headline table, walkthrough, per-file findings by severity, dependency deltas, consistency observations, maintainability observations, skipped files, and linter coverage. Chat gets a short headline plus the path. Per-tool stats in the manifest record `pre_filter` and `post_filter` counts, so the noise removed by diff filtering is visible per tool. ## Telemetry Append-only JSONL at `~/.claude/cache/audit-code/runs.jsonl`. Two record kinds, both written by `scripts/telemetry.py`: `subagent_run` — one row per agent per run: `run_id`, `repo`, `mode`, `agent`, `model`, `input_tokens`, `output_tokens`, `duration_ms`, `finding_count`. `verdict` — one row per finding the user adjudicated: `run_id`, `agent`, `rule_id`, `file`, `line`, `verdict` (one of `kept`, `dismissed`, `false_positive`), and an optional note. Runs are logged by a single call after the fan-out completes: ``` echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \ | uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \ --output-dir --run-id --repo \ --mode --usage-json - ``` `log-run.py` reads token/duration metadata from the piped JSON and reads `finding_count` itself by counting entries in each `findings-.json`. `--log-path` overrides the default log location. Reading it back: ``` uv run scripts/review_stats.py ``` Prints a JSON blob with `runs` (distinct run count), `by_agent`, and `by_rule`. Per agent: cumulative tokens, duration, run count, verdict tallies, `precision` (kept / total adjudicated), and `tokens_per_kept`. Per `/`: verdict tallies and precision. The script takes no arguments and always reads the default log path. Verdict capture is interactive and local-mode only. Ref mode does not collect verdicts, so precision figures reflect local runs alone. ## Development ``` uv run --group dev pytest ``` 197 tests as of writing, covering every adapter against a recorded fixture in `tests/fixtures/`, plus diff filtering, git diff parsing, language detection, manifest serialization, slicing, package diffing, telemetry, `log-run.py`, `review_stats.py`, and the `collect-findings.py` CLI. Tests marked `integration` require the real linters on PATH; the marker is declared in `pyproject.toml`. Layout: ``` SKILL.md orchestration procedure (source of truth) agents/*.md 8 subagent prompts scripts/collect-findings.py CLI entry point scripts/runner.py subprocess wrapper, tool resolution, timeouts scripts/git_diff.py base resolution, changed files + added-line ranges scripts/diff_filter.py findings → changed lines only scripts/language_detect.py extension → language, dep manifests, GHA paths scripts/manifest.py manifest dataclasses / JSON schema scripts/slicing.py per-agent manifest subsets scripts/package_diff.py requirements.txt and package.json deltas scripts/adapters/*.py one parser per linter → Finding[] scripts/rules/ bundled opengrep rules scripts/telemetry.py JSONL append helpers scripts/log-run.py post-fan-out telemetry CLI scripts/review_stats.py telemetry aggregation CLI scripts/install-tools.sh tool installer tests/ pytest suite + recorded tool-output fixtures ``` Adding a linter means: write `scripts/adapters/.py` returning `Finding[]`, record a fixture of its real output, add a test, wire it into the right `_run_*_tools` function in `collect-findings.py`, and add its name to the relevant tool set in `scripts/slicing.py` so a subagent actually receives it.