audit-code install-tools.sh:
- Buffer the opengrep release JSON before grep -m1; curl died with (23)
under pipefail when grep quit early.
- Use ${m}: in the PowerShell block; $m: parsed as a scope-qualified var.
- On Arch, skip paru/yay when pacman -Q shows every package installed,
since --needed still invokes sudo.
- Add --check-only (fast, installs nothing, non-zero naming missing tools)
and --user-only (no system package managers, no sudo).
log-run.py (both skills): put the skill dir on sys.path so running it as
a script from any cwd no longer raises ModuleNotFoundError.
audit-terraform: move deps from requirements.txt into pyproject
dependency groups and add scripts/install-tools.sh (uv sync --group tools,
then check trivy, tflint, tofu, terragrunt, gh).
Both SKILL.md files gain a 0.5 Preflight step and call scripts through
uv run --project ${SKILL_DIR}. tools_unavailable is now a map of tool to
exact install command; audit-terraform skips trivy when absent and stops
with an install hint instead of crashing when tofu/terragrunt is missing.
13 KiB
name, description
| name | description |
|---|---|
| audit-code | Review source-code changes (Python, JavaScript/TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions) by running OSS linters (bandit, ruff, eslint, tsc, mypy, opengrep, gitleaks, pip-audit, osv-scanner, SecurityCodeScan via dotnet, selene, luac, PSScriptAnalyzer, InjectionHunter, actionlint, zizmor), filtering findings to the diff, and dispatching eight parallel review agents for a PR walkthrough plus security triage, type safety, dependencies, consistency, secrets, maintainability, and GitHub Actions review. Two modes — interactive review of the local working tree, or written report for a git ref / GitHub PR. This is the automated tool-driven audit; for a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on phrasing like "audit this code", "run the linters on this change", or "/audit-code". |
audit-code
You review source-code changes by running the OSS SAST stack against the repo, diff-filtering the output, and fanning out eight review subagents in parallel against per-agent manifest slices. One of the eight is a walkthrough agent that produces the reviewer-facing summary of what the PR does.
Always announce at start: "Using audit-code to walk through the change and audit for security, type safety, dependencies, consistency, secrets, and maintainability."
When to invoke
- User says "review the code" / "review this PR" / "review my changes" / etc.
- User runs
/audit-codewith or without an argument.
Modes
| Invocation | Mode | Diff target | Output |
|---|---|---|---|
/audit-code |
local | working tree vs origin | interactive walkthrough in chat |
/audit-code <ref> |
ref | <ref> vs origin |
audit-code-<short>.md + chat |
/audit-code <pr_number> |
ref | PR head ref vs base | same as ref mode |
For ref/PR mode work in a fresh worktree. For local mode work in the user's current repo.
Procedure
0. Resolve SKILL_DIR
${SKILL_DIR} below means the absolute directory containing this SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
0.5 Preflight — every run
bash ${SKILL_DIR}/scripts/install-tools.sh --check-only
It installs nothing and returns in milliseconds when everything is present,
so run it every time. Each plugin version runs from its own directory with a
fresh, empty .venv, so expect it to fail on the first run after an update.
- Exit 0: continue.
- Non-zero: it lists what is missing. Run
bash ${SKILL_DIR}/scripts/install-tools.sh --user-only(Python tools into${SKILL_DIR}/.venv, opengrep, user-scope PowerShell modules; never sudo), then re-run--check-only. - Still missing (
gitleaks,osv-scanner,ghare system packages): ask the user before runningbash ${SKILL_DIR}/scripts/install-tools.shwithout--user-only— it installs through brew/paru/yay/pacman and may call sudo. If they decline, continue: the collection script records each missing tool intools_unavailablewith its install command.
1. Resolve mode, repo identity, and worktree
First resolve NWO (owner/repo) so every gh call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
Always pass --repo "$NWO" on gh calls — do not let gh autodetect from
the cwd, since that runs git internally and fails in jj-only workspaces
with fatal: not a git repository.
Then resolve mode:
- No argument: mode =
local,REPO= cwd. - Argument matches
^[0-9]+$: GitHub PR number. Resolve viagh pr view <PR> --repo "$NWO" --json headRefName,baseRefName. Ifghis missing, stop with: "Installghand rungh auth login, or pass a git ref instead." IfNWOis empty, stop with: "Cannot determine GitHub repo from this directory. Run from inside a checkout of the repo, or pass an explicit ref." - Any other string: git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being a git repo:
-
If a native worktree tool (e.g.
EnterWorktree) is available AND the cwd is a git checkout of$NWO, prefersuperpowers:using-git-worktreesto create~/.claude/cache/audit-code/<short-ref>/. -
Otherwise (jj workspaces, or invoked from outside the repo), do a fresh clone — this never touches the surrounding workspace:
gh repo clone "$NWO" ~/.claude/cache/audit-code/<short-ref>/ git -C ~/.claude/cache/audit-code/<short-ref>/ checkout <ref>For PR mode,
<ref>is the head branch returned bygh pr viewabove.
jj users: Local mode works in jj workspaces that have a colocated
.git(the script reads the working tree but resolves the diff base via git). For non-colocated additionaljj workspace addcheckouts, run ref/PR mode instead — the fresh-clone path works regardless of cwd.
2. Resolve output directory
- Ref mode:
OUTPUT = <REPO>/.audit-code/ - Local mode:
OUTPUT = ~/.claude/cache/audit-code/local-<UTC-timestamp>/
Create the directory.
3. Run the collection script
uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/collect-findings.py \
--repo <REPO> --head <head-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
- If exit code != 0, read
<OUTPUT>/manifest.jsonerrors[]and report them verbatim. Stop. - The script handles the origin-base fetch, language detection, linter dispatch, diff filtering, and writes manifest + per-agent slices.
4. Fan out eight subagents IN PARALLEL
Dispatch eight Task subagents in a single message. Each gets the system
prompt from ${SKILL_DIR}/agents/<agent>.md and these
variables substituted:
MANIFEST = <OUTPUT>/manifest-<agent>.jsonREPO = <REPO>MODE = <local|ref>OUTPUT = <OUTPUT>/findings-<agent>.json(walkthrough uses the same filename but its payload is the walkthrough JSON, not findings).
Agents: walkthrough-reviewer, security-triage-reviewer,
type-safety-reviewer, dependency-reviewer, consistency-reviewer,
secrets-reviewer, maintainability-reviewer, gha-reviewer.
walkthrough-reviewer produces the reviewer-facing PR summary — overview
plus adaptive per-file walkthrough. It emits no findings and does no
triage.
maintainability-reviewer runs on Haiku 4.5 by default — its work is
deterministic-tool triage and doesn't benefit from a larger model.
Override with --maintainability-model.
5. Aggregate findings
After all eight return:
- Load
findings-walkthrough-reviewer.jsonseparately as thewalkthroughpayload (overview+files[]). Discard if malformed and note in summary. - Load each
findings-<agent>.jsonfor the remaining seven. Discard malformed files (note in summary). - Deduplicate findings sharing
{file, line, rule_id}. First-to-finish wins; addalso_flagged_bywith the loser's agent + rule_id. - Group findings by severity: critical, high, medium, low, info.
5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
once. It scans <OUTPUT>/findings-<agent>.json for each agent and
appends one subagent_run row to ~/.claude/cache/audit-code/runs.jsonl
with the finding count read from the file, plus the token / duration
metadata you pass in.
Build a JSON blob from each Task call's <usage> block and pipe it in:
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"security-triage-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"type-safety-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"dependency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"secrets-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"maintainability-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N},
"gha-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | uv run --project ${SKILL_DIR} python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
RUN_ID is a short random hex (8 chars) generated once at the start of
the run; record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not skip it.
Per-agent default models (used for AGENT_MODEL and as defaults if no
override flag is supplied):
| Agent | Default model | Override flag |
|---|---|---|
security-triage-reviewer |
Sonnet | --security-model |
type-safety-reviewer |
Sonnet | --type-safety-model |
dependency-reviewer |
Sonnet | --dependency-model |
consistency-reviewer |
Sonnet | --consistency-model |
secrets-reviewer |
Sonnet | --secrets-model |
maintainability-reviewer |
Haiku 4.5 | --maintainability-model |
walkthrough-reviewer |
Sonnet | --walkthrough-model |
gha-reviewer |
Sonnet | --gha-model |
6. Output
Local mode (interactive)
Send a chat message structured like:
Reviewed N files (X py, Y ts). N findings (n critical, n high, n medium, n low).
Tools: bandit ruff eslint opengrep ... ✓ | mypy ✗ (not on PATH)
WALKTHROUGH
<overview paragraph>
Substantive changes:
- <path> — <one-to-three sentences>
- ...
Trivial:
- <path> — <one-liner>
- ...
CRITICAL (n)
- <file>:<line> (<rule_id> / <CWE>) — <issue>
fix:
<code>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
DEPENDENCY (n)
- <manifest>: <package> <from> → <to> ...
CONSISTENCY (n)
- <file>: <issue>
MAINTAINABILITY (n)
- <file>:<line> — <issue>
Full findings: <OUTPUT>/findings-*.json
Ask me to drill in, expand a section, or generate fixes as patches.
Ref mode (report)
Write <REPO>/audit-code-<short-ref>.md with:
- Headline table — files reviewed, findings by severity, tools used / unavailable
- Walkthrough — overview paragraph, then two subsections:
- "Substantive changes" (each substantive file as a subheading with the summary beneath)
- "Trivial changes" (one-line bullets, collapsible if long)
- Per-file findings grouped by severity, each with
question:framings - Dependency changes with CVE deltas
- Consistency observations (separate section)
- Maintainability observations (dead code, complexity, duplication)
- Skipped files (unsupported languages)
- Linter coverage (which tools ran)
Chat gets a 5-line headline + file path.
7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed the findings, ask for verdicts so the skill
can measure precision over time. This applies to local mode only — prompt
interactively (skip on --no-feedback).
For each finding, capture {kept | dismissed | false_positive} plus an
optional note. Append a verdict record per finding to
~/.claude/cache/audit-code/runs.jsonl via
scripts.telemetry.append_verdict.
Ref-mode reports do not collect verdicts.
review_stats.pytakes no arguments; it only aggregates what local mode already wrote.
8. Always surface the worktree path
In ref mode, end with: "Worktree left at <REPO> for follow-up review."
Failure modes
- No supported source files in diff: manifest's
errors[]populated, script exited non-zero. Surface verbatim. - Unsupported source language(s) in diff: the script keeps reviewing the
supported files but appends an
errors[]entry per language, of the formunsupported language not reviewed: <Language> (<paths>). Surface these prominently at the top of the chat summary / report so the user knows coverage was partial. Example:⚠️ Go file changed but not reviewed: cmd/server.go (Go support not in this skill yet). - No
gh: see step 1. - Tools missing on PATH: non-fatal;
tools_unavailablemaps each to its install command. Agents acknowledge degraded coverage. - PowerShell files changed but no
pwsh:psscriptanalyzerandinjectionhunterreportnot on PATH/<module> not installed. Both come fromscripts/install-tools.sh; InjectionHunter is the injection SAST pass, so its absence means PowerShell injection flaws went unchecked — say so rather than reporting the change as clean. - One agent fails: report the other six and note the agent that failed.