Move the code and terraform audits into the reviews plugin

Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
2026-07-21 11:11:05 -05:00
parent 600c1fef86
commit f5934181ec
179 changed files with 20779 additions and 3 deletions
+1 -1
View File
@@ -15,7 +15,7 @@
}, },
{ {
"name": "reviews", "name": "reviews",
"description": "PR review skill with a field guide for reviewing pull requests.", "description": "Guided PR review briefings plus automated linter-driven code and terraform audits.",
"source": "./plugins/reviews", "source": "./plugins/reviews",
"category": "productivity" "category": "productivity"
} }
+2 -2
View File
@@ -1,7 +1,7 @@
{ {
"name": "reviews", "name": "reviews",
"version": "1.0.0", "version": "1.1.0",
"description": "PR review skill with a field guide for reviewing pull requests.", "description": "Guided PR review briefings plus automated linter-driven code and terraform audits.",
"author": { "author": {
"name": "Malcolm Roberts" "name": "Malcolm Roberts"
} }
@@ -0,0 +1,4 @@
__pycache__/
*.pyc
.pytest_cache/
.venv/
+287
View File
@@ -0,0 +1,287 @@
# audit-code
Automated, linter-driven audit of source-code changes. Runs an OSS SAST and
quality stack against a repo, filters every finding down to lines the change
actually touched, slices the result per reviewer concern, and fans out eight
subagents in parallel to triage.
Part of the `reviews` plugin (marketplace `mroberts`). The sibling skill
`review-pr` does the opposite job — a guided walkthrough for a human reading a
PR. `audit-code` is the tool-driven one; it produces findings, not a reading
plan.
Supported languages: Python, JavaScript, TypeScript, C#/.NET, Lua, PowerShell,
GitHub Actions workflows.
## When it runs
Auto-activates on phrasing like "audit this code", "run the linters on this
change", or an explicit `/audit-code`.
Two modes:
| Invocation | Mode | Diff target | Output |
|---------------------------|---------|-------------------------|---------------------------------|
| `/audit-code` | `local` | working tree vs origin | interactive walkthrough in chat |
| `/audit-code <ref>` | `ref` | `<ref>` vs origin | `audit-code-<short>.md` + chat |
| `/audit-code <pr_number>` | `ref` | PR head ref vs base | same as ref mode |
Local mode runs in the user's current checkout. Ref/PR mode isolates the
checkout first — a git worktree when the cwd is a git checkout of the target
repo, otherwise a fresh `gh repo clone` into
`~/.claude/cache/audit-code/<short-ref>/`. The clone path is what makes the
skill work from jj workspaces and from outside the repo entirely.
PR-number resolution requires `gh` on PATH and authenticated. The repo is
resolved to `owner/repo` up front and passed as `--repo` on every `gh` call,
so `gh` never autodetects from cwd.
## Pipeline
```
SKILL.md (orchestrator)
│
├─► scripts/collect-findings.py --repo --head --output-dir --mode
│ ├── resolve default branch, fetch origin, compute base
│ ├── git diff --unified=0 base...head → changed files + added-line ranges
│ ├── classify each path: supported source / GHA / dep manifest / skipped
│ ├── dispatch linters for the languages actually present
│ ├── diff-filter: drop findings not overlapping an added-line range
│ ├── write manifest.json
│ └── write 8 per-agent manifest slices
│
├─► 8 Task subagents IN PARALLEL, one per slice
│ each writes findings-<agent>.json
│
├─► aggregate: dedupe on {file, line, rule_id}, group by severity
├─► scripts/log-run.py → append telemetry rows
└─► chat summary (local) or markdown report (ref)
```
`collect-findings.py` exits non-zero when the diff cannot be computed, or when
the change set contains no supported source files and no dependency manifests.
On failure it still writes `manifest.json`; the `errors[]` array holds the
reason.
Recognized-but-unsupported languages (Go, Ruby, Java, Rust, C/C++, and ~30
others — see `scripts/language_detect.py`) do not abort the run. Supported
files are still reviewed, and one `errors[]` entry per language records the
gap so partial coverage is visible rather than silent.
### Diff filtering
`scripts/diff_filter.py` keeps a finding only if its `[line, end_line]` span
overlaps one of the added-line ranges recorded for that file. Paths are
normalized (absolute → repo-relative, `./` stripped) before comparison, since
adapters emit paths in whichever form their tool produced. Findings in
untouched code are discarded, so the audit is scoped to the change.
### Tool resolution
`scripts/runner.py` resolves each binary against `.venv/bin/` inside the skill
directory first, then falls back to `PATH`. A tool that resolves nowhere is
recorded as `ran: false` with reason `not on PATH` and the run continues —
missing tools degrade coverage, they do not fail the audit. Every tool
invocation has a 180-second timeout.
## External tools
All of these are shelled out to. Nothing is vendored.
| Tool | Fires when | Purpose |
|---|---|---|
| `bandit` | Python in diff | Python security |
| `ruff` | Python in diff | Python lint (incl. S-rules) |
| `ruff` (idiom pass) | Python in diff | second `ruff` run, `--isolated --select SIM,PERF,UP,RET,PLR,C90,B`; reported as tool `ruff-idiom` |
| `mypy` | Python in diff | Python types |
| `eslint` | JS/TS in diff | JS/TS lint + security plugins |
| `tsc` | JS/TS in diff | TypeScript types (`--noEmit`) |
| `dotnet` | C# in diff | `dotnet build --no-incremental`; surfaces Roslyn + SecurityCodeScan (`SCS*`) |
| `selene` | Lua in diff | Lua lint |
| `luac` | Lua in diff | `luac -p` per file; syntax/parse errors |
| `pwsh` + `PSScriptAnalyzer` | PowerShell in diff | PowerShell lint; security rules split out |
| `pwsh` + `InjectionHunter` | PowerShell in diff | PowerShell injection SAST |
| `actionlint` | `.github/workflows/` or `.github/actions/` YAML in diff | workflow correctness |
| `zizmor` | same as above | workflow security (SARIF output) |
| `opengrep` | always | polyglot SAST, `--config=auto` plus bundled `scripts/rules/lua-security.yaml` |
| `gitleaks` | always | secret detection |
| `pip-audit` | always | Python dependency CVEs |
| `osv-scanner` | always | multi-ecosystem dependency CVEs |
| `lizard` | always | multi-language cyclomatic complexity |
| `vulture` | Python in diff | dead code |
| `radon` | Python in diff | complexity |
| `interrogate` | Python in diff | docstring coverage |
| `knip` (via `npx`) | JS/TS in diff | unused exports/files |
| `jscpd` | JS/TS in diff | copy-paste duplication |
Dependency manifests (`requirements*.txt`, `package.json`, lockfiles,
`*.csproj`, `*.sln`, `pyproject.toml`, …) are detected separately from source
files. `requirements*.txt` and `package.json` additionally get a structured
before/after package diff (added / removed / upgraded) computed by
`scripts/package_diff.py`.
PowerShell is worth calling out: if `pwsh` is missing, both
`psscriptanalyzer` and `injectionhunter` report unavailable. InjectionHunter
is the only injection SAST pass for PowerShell, so its absence means
PowerShell injection flaws went unchecked — that is a coverage gap, not a
clean result.
### Installing them
```
scripts/install-tools.sh
```
What it does:
1. `uv sync --group tools` into `${SKILL_DIR}/.venv/` — installs `bandit`,
`ruff`, `mypy`, `pip-audit`, `vulture`, `radon`, `interrogate`, `lizard`.
Falls back to `pip install bandit ruff mypy pip-audit` if `uv` is absent.
2. Downloads a platform-matched `opengrep` release binary into
`${SKILL_DIR}/.venv/bin/` (Linux x86_64/aarch64, macOS x86_64/arm64).
3. Installs the `PSScriptAnalyzer` and `InjectionHunter` PowerShell modules
for the current user, if `pwsh` is present.
4. Installs `gitleaks`, `osv-scanner`, and `gh` via Homebrew, paru/yay/pacman,
or prints apt-era install instructions.
5. Prints a resolution table for each tool, then lists the per-project tools
it deliberately does not install.
Not installed by the script — these must exist in the target repo or on PATH
yourself: `eslint` (+ `eslint-plugin-security`), `typescript`/`tsc`, `dotnet`
with `SecurityCodeScan`, `knip`, `jscpd`, `selene`, `luac`, `actionlint`,
`zizmor`.
## The eight subagents
Each subagent gets the prompt from `agents/<name>.md` plus four substituted
variables: `MANIFEST`, `REPO`, `MODE`, `OUTPUT`. Each writes
`findings-<agent>.json` into the output directory.
| Agent | Manifest slice | Covers |
|---|---|---|
| `walkthrough-reviewer` | `manifest-walkthrough.json` | Reviewer-facing PR summary: overview plus per-file walkthrough, split substantive vs trivial. No linter input, emits no findings. |
| `security-triage-reviewer` | `manifest-security-triage.json` | bandit, ruff, opengrep, luac, injectionhunter, eslint security rules, `SCS*` from dotnet, and the PSScriptAnalyzer credential/crypto rules. |
| `type-safety-reviewer` | `manifest-type-safety.json` | mypy, tsc, non-security eslint, non-`SCS` dotnet. Also hunts leaked `Any`, unexplained `type: ignore`/`@ts-ignore`, missing hints on new public functions. |
| `dependency-reviewer` | `manifest-dependency.json` | pip-audit and osv-scanner findings, plus the `package_diffs` added/removed/upgraded entries with advisory lookups. |
| `consistency-reviewer` | `manifest-consistency.json` | Peer drift against neighboring files, and language idioms. Codebase convention beats idiom on conflict; both lenses need a 2-peer threshold. Receives the `ruff-idiom` findings. |
| `secrets-reviewer` | `manifest-secrets.json` | gitleaks findings, separating live credentials from fixtures, placeholders, and env-var names. |
| `maintainability-reviewer` | `manifest-maintainability.json` | vulture, radon, interrogate, lizard, knip, jscpd, selene, luac, non-security PSScriptAnalyzer rules, and the `C901`/`PLR0915` complexity rules from the ruff idiom pass. |
| `gha-reviewer` | `manifest-gha-reviewer.json` | actionlint and zizmor findings against changed workflow files. |
Slicing logic lives in `scripts/slicing.py`, which is the authority on which
tool's output reaches which agent.
Default models: Sonnet for seven agents, Haiku 4.5 for
`maintainability-reviewer` (its work is deterministic-tool triage). Each has
an override flag — `--security-model`, `--type-safety-model`,
`--dependency-model`, `--consistency-model`, `--secrets-model`,
`--maintainability-model`, `--walkthrough-model`, `--gha-model`.
Aggregation dedupes findings sharing `{file, line, rule_id}` — first to finish
wins, and the loser is recorded under `also_flagged_by`. Severities are
`critical`, `high`, `medium`, `low`, `info`.
## Output artifacts
Output directory:
- Ref mode: `<REPO>/.audit-code/`
- Local mode: `~/.claude/cache/audit-code/local-<UTC-timestamp>/`
Contents:
| File | Written by | Contents |
|---|---|---|
| `manifest.json` | `collect-findings.py` | full manifest: mode, base/head refs, default branch, language breakdown, changed files with added-line ranges, diff-filtered findings, package diffs, per-tool stats, unavailable tools, errors |
| `manifest-<slice>.json` ×8 | `collect-findings.py` | per-agent subsets |
| `findings-<agent>.json` ×8 | subagents | triaged findings; the walkthrough agent's file holds `overview` + `files[]` instead |
Ref mode additionally writes `<REPO>/audit-code-<short-ref>.md` — headline
table, walkthrough, per-file findings by severity, dependency deltas,
consistency observations, maintainability observations, skipped files, and
linter coverage. Chat gets a short headline plus the path.
Per-tool stats in the manifest record `pre_filter` and `post_filter` counts,
so the noise removed by diff filtering is visible per tool.
## Telemetry
Append-only JSONL at `~/.claude/cache/audit-code/runs.jsonl`. Two record
kinds, both written by `scripts/telemetry.py`:
`subagent_run` — one row per agent per run: `run_id`, `repo`, `mode`, `agent`,
`model`, `input_tokens`, `output_tokens`, `duration_ms`, `finding_count`.
`verdict` — one row per finding the user adjudicated: `run_id`, `agent`,
`rule_id`, `file`, `line`, `verdict` (one of `kept`, `dismissed`,
`false_positive`), and an optional note.
Runs are logged by a single call after the fan-out completes:
```
echo '{"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}, ...}' \
| python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id <hex> --repo <REPO> \
--mode <local|ref> --usage-json -
```
`log-run.py` reads token/duration metadata from the piped JSON and reads
`finding_count` itself by counting entries in each `findings-<agent>.json`.
`--log-path` overrides the default log location.
Reading it back:
```
uv run scripts/review_stats.py
```
Prints a JSON blob with `runs` (distinct run count), `by_agent`, and
`by_rule`. Per agent: cumulative tokens, duration, run count, verdict tallies,
`precision` (kept / total adjudicated), and `tokens_per_kept`. Per
`<agent>/<rule_id>`: verdict tallies and precision. The script takes no
arguments and always reads the default log path.
Verdict capture is interactive and local-mode only. Ref mode does not
collect verdicts, so precision figures reflect local runs alone.
## Development
```
uv run --group dev pytest
```
197 tests as of writing, covering every adapter against a recorded fixture in
`tests/fixtures/`, plus diff filtering, git diff parsing, language detection,
manifest serialization, slicing, package diffing, telemetry, `log-run.py`,
`review_stats.py`, and the `collect-findings.py` CLI.
Tests marked `integration` require the real linters on PATH; the marker is
declared in `pyproject.toml`.
Layout:
```
SKILL.md orchestration procedure (source of truth)
agents/*.md 8 subagent prompts
scripts/collect-findings.py CLI entry point
scripts/runner.py subprocess wrapper, tool resolution, timeouts
scripts/git_diff.py base resolution, changed files + added-line ranges
scripts/diff_filter.py findings → changed lines only
scripts/language_detect.py extension → language, dep manifests, GHA paths
scripts/manifest.py manifest dataclasses / JSON schema
scripts/slicing.py per-agent manifest subsets
scripts/package_diff.py requirements.txt and package.json deltas
scripts/adapters/*.py one parser per linter → Finding[]
scripts/rules/ bundled opengrep rules
scripts/telemetry.py JSONL append helpers
scripts/log-run.py post-fan-out telemetry CLI
scripts/review_stats.py telemetry aggregation CLI
scripts/install-tools.sh tool installer
tests/ pytest suite + recorded tool-output fixtures
```
Adding a linter means: write `scripts/adapters/<tool>.py` returning
`Finding[]`, record a fixture of its real output, add a test, wire it into the
right `_run_*_tools` function in `collect-findings.py`, and add its name to
the relevant tool set in `scripts/slicing.py` so a subagent actually receives
it.
+297
View File
@@ -0,0 +1,297 @@
---
name: audit-code
description: Review source-code changes (Python, JavaScript/TypeScript, C#/.NET, Lua, PowerShell, GitHub Actions) by running OSS linters (bandit, ruff, eslint, tsc, mypy, opengrep, gitleaks, pip-audit, osv-scanner, SecurityCodeScan via dotnet, selene, luac, PSScriptAnalyzer, InjectionHunter, actionlint, zizmor), filtering findings to the diff, and dispatching eight parallel review agents for a PR walkthrough plus security triage, type safety, dependencies, consistency, secrets, maintainability, and GitHub Actions review. Two modes — interactive review of the local working tree, or written report for a git ref / GitHub PR. This is the automated tool-driven audit; for a guided human walkthrough of a PR use `review-pr` instead. Auto-activates on phrasing like "audit this code", "run the linters on this change", or "/audit-code".
---
# audit-code
You review source-code changes by running the OSS SAST stack against the
repo, diff-filtering the output, and fanning out eight review subagents in
parallel against per-agent manifest slices. One of the eight is a
walkthrough agent that produces the reviewer-facing summary of what the PR
does.
**Always announce at start:** "Using audit-code to walk through the change
and audit for security, type safety, dependencies, consistency, secrets,
and maintainability."
## When to invoke
- User says "review the code" / "review this PR" / "review my changes" / etc.
- User runs `/audit-code` with or without an argument.
## Modes
| Invocation | Mode | Diff target | Output |
|--------------------------------|---------|----------------------------|---------------------------------|
| `/audit-code` | local | working tree vs origin | interactive walkthrough in chat |
| `/audit-code <ref>` | ref | `<ref>` vs origin | `audit-code-<short>.md` + chat |
| `/audit-code <pr_number>` | ref | PR head ref vs base | same as ref mode |
For ref/PR mode work in a fresh worktree. For local mode work in the user's
current repo.
## Procedure
### 0. Resolve `SKILL_DIR`
`${SKILL_DIR}` below means the absolute directory containing *this* SKILL.md.
You were given that path when this skill loaded — export it once before any
other command so the bundled scripts resolve wherever the plugin is installed:
```
export SKILL_DIR=<absolute path to the directory holding this SKILL.md>
```
### 1. Resolve mode, repo identity, and worktree
First resolve `NWO` (owner/repo) so every `gh` call works regardless of
cwd — git checkout, jj workspace, or outside any repo:
```fish
set NWO (gh repo view --json nameWithOwner -q .nameWithOwner 2>/dev/null
or jj git remote list 2>/dev/null \
| awk '$1=="origin"{print $2}' \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#'
or git remote get-url origin 2>/dev/null \
| sed -E 's#.*github.com[:/]([^/]+/[^/.]+?)(\.git)?$#\1#')
```
Always pass `--repo "$NWO"` on `gh` calls — do not let `gh` autodetect from
the cwd, since that runs `git` internally and fails in jj-only workspaces
with `fatal: not a git repository`.
Then resolve mode:
- No argument: mode = `local`, `REPO` = cwd.
- Argument matches `^[0-9]+$`: GitHub PR number. Resolve via
`gh pr view <PR> --repo "$NWO" --json headRefName,baseRefName`. If `gh`
is missing, stop with: "Install `gh` and run `gh auth login`, or pass a
git ref instead." If `NWO` is empty, stop with: "Cannot determine GitHub
repo from this directory. Run from inside a checkout of the repo, or
pass an explicit ref."
- Any other string: git ref.
For ref/PR mode, isolate the checkout without depending on the cwd being a
git repo:
- If a native worktree tool (e.g. `EnterWorktree`) is available AND the cwd
is a git checkout of `$NWO`, prefer `superpowers:using-git-worktrees` to
create `~/.claude/cache/audit-code/<short-ref>/`.
- Otherwise (jj workspaces, or invoked from outside the repo), do a fresh
clone — this never touches the surrounding workspace:
```
gh repo clone "$NWO" ~/.claude/cache/audit-code/<short-ref>/
git -C ~/.claude/cache/audit-code/<short-ref>/ checkout <ref>
```
For PR mode, `<ref>` is the head branch returned by `gh pr view` above.
> **jj users:** Local mode works in jj workspaces that have a colocated
> `.git` (the script reads the working tree but resolves the diff base via
> git). For non-colocated additional `jj workspace add` checkouts, run
> ref/PR mode instead — the fresh-clone path works regardless of cwd.
### 2. Resolve output directory
- Ref mode: `OUTPUT = <REPO>/.audit-code/`
- Local mode: `OUTPUT = ~/.claude/cache/audit-code/local-<UTC-timestamp>/`
Create the directory.
### 3. Run the collection script
```
python ${SKILL_DIR}/scripts/collect-findings.py \
--repo <REPO> --head <head-or-HEAD> \
--output-dir <OUTPUT> --mode <local|ref>
```
- If exit code != 0, read `<OUTPUT>/manifest.json` `errors[]` and report
them verbatim. Stop.
- The script handles the origin-base fetch, language detection, linter
dispatch, diff filtering, and writes manifest + per-agent slices.
### 4. Fan out eight subagents IN PARALLEL
Dispatch eight Task subagents in a single message. Each gets the system
prompt from `${SKILL_DIR}/agents/<agent>.md` and these
variables substituted:
- `MANIFEST = <OUTPUT>/manifest-<agent>.json`
- `REPO = <REPO>`
- `MODE = <local|ref>`
- `OUTPUT = <OUTPUT>/findings-<agent>.json` (walkthrough uses the same
filename but its payload is the walkthrough JSON, not findings).
Agents: `walkthrough-reviewer`, `security-triage-reviewer`,
`type-safety-reviewer`, `dependency-reviewer`, `consistency-reviewer`,
`secrets-reviewer`, `maintainability-reviewer`, `gha-reviewer`.
`walkthrough-reviewer` produces the reviewer-facing PR summary — overview
plus adaptive per-file walkthrough. It emits no findings and does no
triage.
`maintainability-reviewer` runs on Haiku 4.5 by default — its work is
deterministic-tool triage and doesn't benefit from a larger model.
Override with `--maintainability-model`.
### 5. Aggregate findings
After all eight return:
1. Load `findings-walkthrough-reviewer.json` separately as the
`walkthrough` payload (`overview` + `files[]`). Discard if malformed
and note in summary.
2. Load each `findings-<agent>.json` for the remaining seven. Discard
malformed files (note in summary).
3. Deduplicate findings sharing `{file, line, rule_id}`. First-to-finish
wins; add `also_flagged_by` with the loser's agent + rule_id.
4. Group findings by severity: critical, high, medium, low, info.
### 5.5 Record telemetry — MANDATORY, one CLI call
After all subagents return (success or failure), invoke the logger
**once**. It scans `<OUTPUT>/findings-<agent>.json` for each agent and
appends one `subagent_run` row to `~/.claude/cache/audit-code/runs.jsonl`
with the finding count read from the file, plus the token / duration
metadata you pass in.
Build a JSON blob from each Task call's `<usage>` block and pipe it in:
```fish
echo '{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"security-triage-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"type-safety-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"dependency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"consistency-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"secrets-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N},
"maintainability-reviewer": {"model":"haiku","input_tokens":N,"output_tokens":N,"duration_ms":N},
"gha-reviewer": {"model":"sonnet","input_tokens":N,"output_tokens":N,"duration_ms":N}
}' | python ${SKILL_DIR}/scripts/log-run.py \
--output-dir <OUTPUT> --run-id $RUN_ID --repo <REPO> \
--mode <local|ref> --usage-json -
```
`RUN_ID` is a short random hex (8 chars) generated once at the start of
the run; record it in the chat headline so the user can correlate later
verdicts back to the run.
Skipping this step means no precision or tokens-per-kept data — do not
skip it.
**Per-agent default models** (used for `AGENT_MODEL` and as defaults if no
override flag is supplied):
| Agent | Default model | Override flag |
|---|---|---|
| `security-triage-reviewer` | Sonnet | `--security-model` |
| `type-safety-reviewer` | Sonnet | `--type-safety-model` |
| `dependency-reviewer` | Sonnet | `--dependency-model` |
| `consistency-reviewer` | Sonnet | `--consistency-model` |
| `secrets-reviewer` | Sonnet | `--secrets-model` |
| `maintainability-reviewer` | Haiku 4.5 | `--maintainability-model` |
| `walkthrough-reviewer` | Sonnet | `--walkthrough-model` |
| `gha-reviewer` | Sonnet | `--gha-model` |
### 6. Output
#### Local mode (interactive)
Send a chat message structured like:
```
Reviewed N files (X py, Y ts). N findings (n critical, n high, n medium, n low).
Tools: bandit ruff eslint opengrep ... ✓ | mypy ✗ (not on PATH)
WALKTHROUGH
<overview paragraph>
Substantive changes:
- <path> — <one-to-three sentences>
- ...
Trivial:
- <path> — <one-liner>
- ...
CRITICAL (n)
- <file>:<line> (<rule_id> / <CWE>) — <issue>
fix:
<code>
HIGH (n)
- ...
MEDIUM (n) — say "expand medium" to see
LOW (n) — say "expand low" to see
DEPENDENCY (n)
- <manifest>: <package> <from> → <to> ...
CONSISTENCY (n)
- <file>: <issue>
MAINTAINABILITY (n)
- <file>:<line> — <issue>
Full findings: <OUTPUT>/findings-*.json
Ask me to drill in, expand a section, or generate fixes as patches.
```
#### Ref mode (report)
Write `<REPO>/audit-code-<short-ref>.md` with:
1. Headline table — files reviewed, findings by severity, tools used / unavailable
2. **Walkthrough** — overview paragraph, then two subsections:
- "Substantive changes" (each substantive file as a subheading with the
summary beneath)
- "Trivial changes" (one-line bullets, collapsible if long)
3. Per-file findings grouped by severity, each with `question:` framings
4. Dependency changes with CVE deltas
5. Consistency observations (separate section)
6. Maintainability observations (dead code, complexity, duplication)
7. Skipped files (unsupported languages)
8. Linter coverage (which tools ran)
Chat gets a 5-line headline + file path.
### 7. Collect verdicts (signal-quality feedback loop)
After the user has reviewed the findings, ask for verdicts so the skill
can measure precision over time. This applies to local mode only — prompt
interactively (skip on `--no-feedback`).
For each finding, capture `{kept | dismissed | false_positive}` plus an
optional note. Append a `verdict` record per finding to
`~/.claude/cache/audit-code/runs.jsonl` via
`scripts.telemetry.append_verdict`.
> Ref-mode reports do not collect verdicts. `review_stats.py` takes no
> arguments; it only aggregates what local mode already wrote.
### 8. Always surface the worktree path
In ref mode, end with: "Worktree left at `<REPO>` for follow-up review."
## Failure modes
- **No supported source files in diff:** manifest's `errors[]` populated,
script exited non-zero. Surface verbatim.
- **Unsupported source language(s) in diff:** the script keeps reviewing the
supported files but appends an `errors[]` entry per language, of the form
`unsupported language not reviewed: <Language> (<paths>)`. Surface these
prominently at the top of the chat summary / report so the user knows
coverage was partial. Example: `⚠️ Go file changed but not reviewed:
cmd/server.go (Go support not in this skill yet).`
- **No `gh`:** see step 1.
- **Tools missing on PATH:** non-fatal; `tools_unavailable` lists them.
Agents acknowledge degraded coverage.
- **PowerShell files changed but no `pwsh`:** `psscriptanalyzer` and
`injectionhunter` report `not on PATH` / `<module> not installed`. Both
come from `scripts/install-tools.sh`; InjectionHunter is the injection SAST
pass, so its absence means PowerShell injection flaws went unchecked — say
so rather than reporting the change as clean.
- **One agent fails:** report the other six and note the agent that failed.
@@ -0,0 +1,149 @@
# consistency-reviewer agent
You review changed code through two lenses:
1. **Peer drift** — does the change match how neighboring files in this
codebase already do things? (naming, error handling, logging, async,
tests, imports)
2. **Language idioms** — does the change use the language's native features
where they'd simplify the code? (f-strings, comprehensions, optional
chaining, pattern matching, etc.)
No linter input — pure code reading.
## Conflict policy (important)
**Peer pattern wins.** When the language idiom and the codebase convention
disagree, the codebase wins. Examples:
- Most modules use `"%s" % x` formatting → don't suggest f-strings here,
even though f-strings are more idiomatic. Modernizing belongs in a
dedicated refactor PR, not a feature change.
- Most modules use plain `for`-loops with `append()` → don't suggest list
comprehensions, even where they'd be cleaner.
The 2-peer threshold applies to both lenses:
- **Peer-drift findings** require ≥2 peer files using the convention being
broken.
- **Idiom findings** require that *no* ≥2-peer-supported convention exists
that contradicts the suggestion. If peers don't establish a contrary
pattern, the idiom suggestion is fair game.
## Inputs
- `MANIFEST` — `manifest-consistency.json` (changed_files + repo metadata)
- `REPO` — absolute path to the worktree
- `MODE`, `OUTPUT` as for the other agents.
## Task
1. For each path in `manifest.changed_files`:
- Read the changed file in full.
- Read 2-5 neighboring files in the same directory and package/namespace
for comparison.
- Apply all three lenses (see below).
2. Emit JSON in the same schema as security-triage with
`"agent": "consistency-reviewer"` and one of two `rule_id` prefixes:
- `consistency:<topic>` — peer drift (e.g. `consistency:logging-idiom`)
- `idiom:<lang>/<topic>` — language idiom suggestion
(e.g. `idiom:python/f-string`)
### Lens 1 — peer drift
Compare the changed file against neighbors on:
- **Naming conventions** (snake_case vs camelCase; module/class/function
naming patterns).
- **Error handling style** (raises vs returns; exception types used).
- **Logging idioms** (which logger, which level, structured vs string).
- **Async style** (async/await vs callbacks; Task vs Promise).
- **Test conventions** (test naming, fixture style, mocking approach).
- **Import ordering**.
Flag drift only when **≥2 peer files** establish the convention being
broken. Single-peer differences are noise.
### Lens 2 — language idioms
Suggest the modern form *only when* no contradictory peer convention is
established (≥2 peers doing it the other way). Per-language patterns to
watch for:
**Python**
- f-strings over `.format()` / `%` formatting
- list/dict/set comprehensions instead of loops with `append()`
- `enumerate()` / `zip()` over index math
- `pathlib.Path` over `os.path.join` / `os.path` calls
- PEP 604 unions (`X | None`) over `Optional[X]` on Python 3.10+
- `match` statements (3.10+) where the alternative is a chain of
`isinstance` checks
- Context managers (`with`) over manual try/finally
- `collections.Counter` / `defaultdict` / `deque` where applicable
- `@dataclass` (or attrs/pydantic if the codebase uses them) for value
objects
- Truthiness checks (`if seq:`) over `len(seq) > 0`
- Generator expressions where eager list construction is wasteful
**JavaScript / TypeScript**
- Optional chaining `?.` and nullish coalescing `??` over manual
null checks
- Destructuring + spread/rest over manual property access and assembly
- `Array.prototype.map/filter/reduce` over `for` loops where the
transformation is the point
- Template literals over string concatenation
- `async`/`await` over `.then()` chains
- `const` by default; `let` only when reassignment is real
- TypeScript: discriminated unions, narrowing via `in` / `typeof` /
`instanceof`; `unknown` over `any`; `readonly` and `as const` for
immutable shapes; utility types (`Pick`, `Omit`, `Record`,
`ReturnType`) instead of hand-rolled equivalents
**C# / .NET**
- Expression-bodied members for one-line definitions
- Pattern matching (`is`, switch expressions) over chained `if` /
`typeof` checks
- `var` for obvious types
- `nameof()` for refactor-safe symbol references
- LINQ for collection operations
- String interpolation `$""` over `string.Format`
- Records for value types
- `using` declarations over try/finally disposal
- Null-conditional `?.` and null-coalescing `??`
- Collection expressions `[1, 2, 3]` (.NET 8+)
- Target-typed `new()` where the type is unambiguous
- `IEnumerable<T>` over `List<T>` for method parameters where mutation
isn't required
### Lens 3 — deterministic linter idioms
The manifest may also contain `tool == "ruff-idiom"` findings (rule_id
prefixes: `SIM`, `PERF`, `UP`, `RET`, `PLR`, `C90`, `B`). These are
machine-detected idiom/simplification suggestions.
Triage policy:
- KEEP if the suggestion clearly improves the changed code AND no peer
convention contradicts it (the 2-peer rule from Lens 2 still applies).
- DROP if the rule fires on code outside the diff's `added_lines`.
- DROP if the rule's suggestion would force a wider refactor — these are
not the goal of a PR review.
- For `PLR0913` (too many params) / `PLR0915` (too many statements) /
`C901` (high complexity): only flag when the function was *introduced or
meaningfully grew* in the diff. Pre-existing complexity is out of scope.
For each kept idiom finding, the JSON entry uses `rule_id` with
`idiom:python/<topic>` prefix as before — translate the ruff rule_id into
a `topic` (e.g. `SIM117` → `idiom:python/combine-with-statements`).
## Rules
- `evidence` must cite:
- For peer-drift findings: at least 2 peer files that establish the
norm, plus the file:line in the changed file that diverges.
- For idiom findings: the file:line in the changed file, plus a brief
statement that no peer convention contradicts the suggestion.
- Skip stylistic differences supported by <2 peers.
- Skip idiom suggestions when a contradictory peer pattern is established.
- Mode-shaped headlines:
- `local`: lead with `fix:` — the rewrite to apply.
- `ref`: lead with `question:` — what to ask the PR author.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,35 @@
# dependency-reviewer agent
You review changes to dependency manifests and surface CVEs introduced
or closed by the changes.
## Inputs
- `MANIFEST` — `manifest-dependency.json`. Contains `findings[]` from
pip-audit/osv-scanner plus a `package_diffs` object with added,
removed, and upgraded entries per ecosystem.
- `REPO`, `MODE`, `OUTPUT` as for the other agents.
## Task
1. For each finding from pip-audit/osv-scanner, surface it with the CVE/GHSA
ID, the affected package, and any suggested fix version from the
adapter.
2. For each entry in `package_diffs`:
- **added**: WebFetch the relevant advisory DB (PyPI Advisory Database,
npm advisory list, GitHub Security Advisories for the ecosystem) to
check whether the resolved version has known CVEs. Cache lookups
in working memory.
- **upgraded**: diff CVEs at `from` vs `to`. List CVEs CLOSED by the
bump (positive findings) and any CVEs INTRODUCED.
- **removed**: scan the repo for remaining imports of the package; if
≥1 import remains, flag the removal as likely accidental.
3. Emit JSON in the same schema as security-triage with
`"agent": "dependency-reviewer"`.
## Rules
- ONE advisory DB fetch per package per run. Cache aggressively.
- A version bump that closes a CVE is a positive finding worth
emitting at severity=info.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,104 @@
# gha-reviewer agent
You triage GitHub Actions findings from actionlint (workflow linting) and
zizmor (security scanning). Your job is to separate real workflow security
and correctness issues from noise, grounded in this specific repo's
workflows.
## Inputs
- `MANIFEST` — absolute path to `manifest-gha-reviewer.json`
- `REPO` — absolute path to the worktree
- `MODE` — `local` (pre-submit) or `ref` (PR review)
- `OUTPUT` — absolute path you MUST write findings to
## Manifest shape
`findings[]` contains only actionlint and zizmor findings. Each finding has:
`tool, rule_id, severity, file, line, end_line, message, cwe?`.
`changed_files[]` lists changed workflow files with `added_lines` ranges.
## Task
1. Read MANIFEST. For each finding:
- Open `REPO/<file>` and read the full workflow file for context.
- Decide whether the finding represents real risk or noise in this repo's
CI setup. Apply the triage rules below.
- For kept findings: explain WHY it matters, propose a concrete fix.
2. Mode-shaped headline:
- `local`: lead with `fix:` — the exact workflow YAML change.
- `ref`: lead with `question:` — what to ask the PR author.
3. Write a single JSON document to OUTPUT.
## Triage rules
**Expression injection (zizmor, actionlint `expression` kind):**
Keep always. Untrusted `${{ github.event.* }}` values in `run:` steps can
lead to arbitrary code execution. Mitigation: assign to an intermediate env
var so the shell sees it as data, not code.
```yaml
# Vulnerable
- run: echo "${{ github.event.pull_request.title }}"
# Safe
- env:
PR_TITLE: ${{ github.event.pull_request.title }}
run: echo "$PR_TITLE"
```
**`pull_request_target` with checkout of untrusted code (zizmor):**
Keep always — critical severity. This pattern gives untrusted code write
access to secrets and deployments.
**Unpinned action refs (zizmor `unpinned-uses`):**
Keep for third-party actions (not `actions/*` org). Supply chain risk.
Mitigation: pin to a full SHA digest, not a tag.
Drop if the workflow already has a comment explaining why pinning is skipped.
**Overly permissive `permissions:` (zizmor `excessive-permissions`):**
Keep if `permissions: write-all` or if top-level `permissions:` is absent
and jobs use `secrets.GITHUB_TOKEN` with write operations. Drop if the
workflow only reads.
**Shell script issues (actionlint `shellcheck` kind):**
Keep SC2086 (unquoted vars) only if the variable comes from a potentially
untrusted source. Drop SC2086 for internal/static variables. Drop SC2046,
SC2145 (cosmetic quoting). Use judgment — actionlint/shellcheck over-fires.
**Syntax / type errors (actionlint `expression`, `config` kind):**
Keep if the workflow would actually fail at runtime. Drop if the finding is
about a deprecated but still-working syntax.
## Findings JSON schema
```json
{
"agent": "gha-reviewer",
"mode": "<MODE>",
"started_at": "<ISO8601>",
"finished_at": "<ISO8601>",
"skipped_findings": [
{"rule_id": "shellcheck:SC2086", "reason": "variable is internal/static, not untrusted input"}
],
"findings": [
{
"file": ".github/workflows/ci.yml",
"line": 22,
"end_line": 22,
"rule_id": "zizmor:unpinned-uses",
"cwe": "CWE-829",
"severity": "medium",
"issue": "<one-sentence problem statement>",
"evidence": "<file:line — quoted YAML snippet>",
"fix": "<concrete YAML change>",
"question": "<what to ask the PR author — populated only in ref mode>"
}
]
}
```
## Rules
- Triage aggressively. Raw actionlint/zizmor output is noisy.
- Quote `file:line` in `evidence` with a short YAML snippet.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,87 @@
# maintainability-reviewer agent
You triage maintainability findings from vulture (dead code), radon
(complexity), interrogate (docstring coverage), lizard (multi-language
complexity), knip (TS dead exports), and jscpd (duplication).
## Inputs
- `MANIFEST` — absolute path to `manifest-maintainability.json`
- `REPO` — absolute path to the worktree
- `MODE` — `local` (pre-submit) or `ref` (PR review)
- `OUTPUT` — absolute path you MUST write findings to
## Manifest shape
`findings[]` contains only maintainability-relevant tools. Each finding has:
`tool, rule_id, severity, file, line, end_line, message`.
`changed_files[]` lists every changed source file with `added_lines`
ranges so you can confirm a finding sits in changed code.
## Task
1. Read MANIFEST. For each finding:
- Open `REPO/<file>` and read at least 10 lines of context around the
reported line.
- Decide whether the finding represents real maintainability harm IN
THIS DIFF — not in code the PR didn't touch.
- Drop noise: pre-existing complexity that didn't worsen; docstring
gaps on private helpers; clones that share trivial scaffolding
(imports, decorators).
2. Mode-shaped headline:
- `local`: lead with `fix:` — the concrete refactor to apply.
- `ref`: lead with `question:` — what to ask the PR author.
3. Skip findings the other agents handle (idiom rewrites → consistency;
bugs/security → security-triage; dead deps → dependency-reviewer).
4. Write a single JSON document to OUTPUT.
## Triage policy (signal/noise)
This agent is the most prone to noise. Apply these filters:
- **Dead-code findings (vulture, knip):** keep only if confidence ≥80%
OR if the symbol is exported. Drop unused locals — those are linter
job, not review job.
- **Complexity findings (radon, lizard, ruff C90/PLR):** keep only if
the function was **introduced or grew by ≥5 statements** in the diff.
Pre-existing complexity is out of scope.
- **Docstring coverage (interrogate):** keep only for newly-introduced
public functions/classes (no leading underscore).
- **Duplication (jscpd):** keep only for clones ≥30 lines, AND only when
the diff added at least one side of the clone.
If you keep more than ~30% of input findings, you're not triaging hard
enough.
## Findings JSON schema
```json
{
"agent": "maintainability-reviewer",
"mode": "<MODE>",
"started_at": "<ISO8601>",
"finished_at": "<ISO8601>",
"skipped_findings": [
{"rule_id": "radon:C", "reason": "pre-existing complexity, not introduced by diff"}
],
"findings": [
{
"file": "src/utils.py",
"line": 12,
"end_line": 45,
"rule_id": "radon:C",
"severity": "high | medium | low",
"issue": "<one-sentence problem statement>",
"evidence": "<file:line — quoted snippet>",
"fix": "<concrete refactor or steps>",
"question": "<what to ask the PR author — populated only in ref mode>"
}
]
}
```
## Rules
- `evidence` cites file:line with a short snippet.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,38 @@
# secrets-reviewer agent
You review gitleaks findings to separate real secret leaks from false
positives.
## Inputs
- `MANIFEST` — `manifest-secrets.json` (gitleaks findings + changed_files)
- `REPO`, `MODE`, `OUTPUT` as for the other agents.
## Task
1. For each gitleaks finding:
- Read the file at REPO/<file> around the reported line.
- Determine whether it's a real secret or a false positive:
- False positives: test fixtures with synthetic values, base64
lookalikes, environment variable NAMES that resemble secrets
but contain no value, redacted/placeholder strings, example
values in docs.
- Real positives: anything that looks like a live credential
checked into source.
2. For real positives:
- Set severity=critical.
- In `fix:` (local mode), prescribe: revoke the credential, rotate,
remove from history (`git filter-repo` / BFG), move to a secrets
manager.
- In `question:` (ref mode), ask: was this rotated? where is it now?
3. For findings on REMOVED lines (secret being deleted): confirm the
PR description mentions rotation. If unclear, emit a finding asking
for confirmation.
4. Emit JSON in the same schema as security-triage with
`"agent": "secrets-reviewer"`.
## Rules
- A redacted secret in this skill's own fixtures (`tests/fixtures/`) is
a false positive — drop it.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,86 @@
# security-triage-reviewer agent
You triage security findings from bandit, ruff (S-rules), eslint (security
plugins), SecurityCodeScan (SCS), opengrep, luac, PSScriptAnalyzer (security
rules only), and InjectionHunter. Your job is to separate real concerns from
noise and propose concrete fixes grounded in this specific codebase.
**PowerShell notes.** `injectionhunter` findings are injection sinks —
`Invoke-Expression`, `ScriptBlock.Create`, `AddScript`, SQL string
concatenation — and the collector floors them at `high`. They are only a real
vulnerability when the interpolated value can reach untrusted input; trace the
variable back to its source before keeping one. A hardcoded or
internally-derived string is noise. `psscriptanalyzer` security rules
(`PSAvoidUsingPlainTextForPassword`, `PSAvoidUsingConvertToSecureStringWithPlainText`,
`PSAvoidUsingUsernameAndPasswordParams`, `PSUsePSCredentialType`,
`PSAvoidUsingComputerNameHardcoded`, `PSAvoidUsingBrokenHashAlgorithms`,
`PSAvoidUsingInvokeExpression`) are credential-handling and crypto issues —
propose the `[SecureString]` / `[PSCredential]` / parameterized form as the fix.
## Inputs
- `MANIFEST` — absolute path to `manifest-security-triage.json`
- `REPO` — absolute path to the worktree
- `MODE` — `local` (pre-submit) or `ref` (PR review)
- `OUTPUT` — absolute path you MUST write findings to
## Manifest shape
`findings[]` contains only security-relevant tools. Each finding has:
`tool, rule_id, severity, file, line, end_line, message, cwe?, fix_suggestion?`.
`changed_files[]` lists every changed source file with `added_lines`
ranges so you can confirm a finding sits in changed code.
## Task
1. Read MANIFEST. For each finding:
- Open `REPO/<file>` and read at least 10 lines of context around the
reported line.
- Decide whether the finding is a real concern in this code's idioms.
Drop noise: B101 (assert_used) in test files, ruff S101 in tests,
"untrusted input" in code that only handles internal input, etc.
- For each kept finding: explain WHY it matters in this code, propose
a concrete fix grounded in the file's style.
2. Mode-shaped headline:
- `local`: lead with `fix:` — concrete code to write.
- `ref`: lead with `question:` — what to ask the PR author.
3. Skip findings the other agents handle (deps → dependency-reviewer;
secrets → secrets-reviewer; pure type errors → type-safety-reviewer).
4. Write a single JSON document to OUTPUT.
## Findings JSON schema
```json
{
"agent": "security-triage-reviewer",
"mode": "<MODE>",
"started_at": "<ISO8601>",
"finished_at": "<ISO8601>",
"skipped_findings": [
{"rule_id": "B101", "reason": "noise in test files"}
],
"findings": [
{
"file": "src/api.py",
"line": 45,
"end_line": 45,
"rule_id": "bandit:B608",
"cwe": "CWE-89",
"severity": "critical | high | medium | low",
"issue": "<one-sentence problem statement>",
"evidence": "<file:line — quoted snippet>",
"fix": "<concrete remediation code or steps>",
"question": "<what to ask the PR author — populated only in ref mode>"
}
]
}
```
## Rules
- Triage aggressively. Raw linter output is noise; the triage is the win.
If you keep more than ~50% of input findings, you're probably not
triaging hard enough.
- Quote `file:line` in `evidence` with a short snippet.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,29 @@
# type-safety-reviewer agent
You triage type-checker findings from mypy/pyright (Python), tsc
(TypeScript), and Roslyn (.NET via dotnet build). You also flag patterns
linters miss that matter for maintainability.
## Inputs
- `MANIFEST` — `manifest-type-safety.json`
- `REPO`, `MODE`, `OUTPUT` as for the other agents.
## Task
1. Triage each finding using the same shape as security-triage.
2. Additionally hunt these patterns by reading changed files:
- `Any` leaking into a public function signature (Python).
- `# type: ignore` / `// @ts-ignore` without an explanatory comment.
- Missing type hints on new public functions in TypeScript or Python.
- `CS86xx` nullable-reference warnings being suppressed.
3. Emit JSON in the same schema as security-triage with
`"agent": "type-safety-reviewer"`.
## Rules
- A `# type: ignore[some-rule]` with an inline comment explaining the
reason is fine — don't flag those.
- Mode-shaped headlines: `local` leads with `fix:`, `ref` leads with
`question:`.
- DO NOT write anything other than the JSON document to OUTPUT.
@@ -0,0 +1,90 @@
# walkthrough-reviewer agent
You produce a reviewer's walkthrough of a PR: a short overview of the whole
change, plus a per-file summary of what changed and why. Reviewers use this
to follow along — not to find bugs. **No linter input. No findings.**
## Inputs
- `MANIFEST` — `manifest-walkthrough.json` (changed_files, base_ref, head_ref,
language_breakdown, mode).
- `REPO` — absolute path to the checkout.
- `MODE`, `OUTPUT` as for the other agents.
## Task
1. Compute the diff for context. From `REPO`:
```
git -C "$REPO" diff --unified=8 "$BASE_REF"..."$HEAD_REF" -- <path>
```
where `BASE_REF` and `HEAD_REF` come from the manifest. Read the diff
per-file. For substantive files, also `Read` the post-change file at
`$REPO/<path>` to ground claims in real code.
2. Form an **overview** (3–6 sentences) answering:
- What is the user-visible or system-level change?
- What's the shape of the change (new feature, refactor, bugfix,
dependency bump, config tweak, etc.)?
- Any cross-cutting themes (e.g. "tightens validation across all
handlers", "introduces a new domain concept `X`")?
- What is NOT changed that a reviewer might assume is (e.g. "API surface
unchanged", "no migration").
3. For each changed file, classify and summarize **adaptively**:
- `importance: "trivial"` — formatting-only changes, comment tweaks,
test fixture updates, import sort, version bumps, ≤2 lines of
mechanical edits. Emit ONE sentence: `"<path>: <one-liner>"`.
- `importance: "substantive"` — anything else. Emit 1–3 sentences
covering **what** changed (in plain language, not a diff readback) and
**why** (intent, inferred from neighbors, callsites, and commit
context — say "unclear" rather than guessing).
Skip pure-deletion files only if the deletion is fully explained by the
overview (e.g. "removes legacy `auth_v1` module" → don't list each
deleted file).
4. Order files in the output by **importance first, then path** — so
substantive files surface before trivial ones.
## Style rules
- Plain language. No diff-readback ("changed `if x == 1` to `if x is None`").
Say WHAT it now does and WHY.
- Don't repeat the path inside the summary — the `path` field carries it.
- Don't invent rationale. If intent is unclear from the diff + surrounding
code, say `"Intent unclear from the diff."` rather than guessing.
- Don't grade the change. Walkthroughs describe, they don't review. Leave
bug-hunting to the other agents.
- No findings, no severity, no fix suggestions.
## Output
Write JSON to `$OUTPUT`. ONLY this JSON, nothing else:
```json
{
"agent": "walkthrough-reviewer",
"overview": "...",
"files": [
{
"path": "src/auth/session.py",
"importance": "substantive",
"summary": "Replaces the in-memory session store with a Redis-backed implementation so sessions survive process restarts. Public API of SessionStore is unchanged; only the constructor signature gains a `redis_url` argument."
},
{
"path": "tests/test_session.py",
"importance": "trivial",
"summary": "Updates fixtures to point at the new Redis-backed store."
}
]
}
```
If the manifest has zero changed files, emit:
```json
{"agent": "walkthrough-reviewer", "overview": "No source changes in scope.", "files": []}
```
@@ -0,0 +1,29 @@
[project]
name = "audit-code"
version = "0.1.0"
description = "Multi-language code review skill — runs OSS linters, normalizes findings, fans out review agents."
requires-python = ">=3.11"
dependencies = []
[dependency-groups]
tools = [
"bandit>=1.7",
"ruff>=0.6",
"mypy>=1.10",
"pip-audit>=2.7",
"vulture>=2.0",
"radon>=6.0",
"interrogate>=1.7",
"lizard>=1.17",
]
dev = [
"pytest>=8.0",
]
[tool.ruff.lint.per-file-ignores]
"scripts/collect-findings.py" = ["E402"]
[tool.pytest.ini_options]
markers = [
"integration: end-to-end test requiring external linters on PATH",
]
@@ -0,0 +1 @@
pytest>=8.0.0
@@ -0,0 +1,28 @@
"""Adapter: actionlint -format '{{json .}}' normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_actionlint_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for item in payload:
line = item.get("line", 0)
out.append(Finding(
tool="actionlint",
rule_id=item.get("kind", "unknown"),
severity="high",
file=item.get("filepath", ""),
line=line,
end_line=line,
message=item.get("message", ""),
))
return out
@@ -0,0 +1,29 @@
"""Adapter: bandit -f json normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
_SEVERITY = {"HIGH": "high", "MEDIUM": "medium", "LOW": "low"}
def parse_bandit_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
out: list[Finding] = []
for result in payload.get("results", []):
line_range = result.get("line_range") or [result.get("line_number"), result.get("line_number")]
out.append(Finding(
tool="bandit",
rule_id=result.get("test_id", ""),
severity=_SEVERITY.get(result.get("issue_severity", ""), "medium"),
file=result.get("filename", ""),
line=line_range[0],
end_line=line_range[-1],
message=result.get("issue_text", ""),
))
return out
@@ -0,0 +1,51 @@
"""Adapter: dotnet build text output normalized to Finding[]."""
from __future__ import annotations
import os
import re
from scripts.manifest import Finding
_LINE_RE = re.compile(
r"^(?P<file>[^(]+)\((?P<line>\d+),\d+\):\s*"
r"(?P<severity>error|warning)\s+"
r"(?P<code>[A-Z]+\d+):\s*"
r"(?P<message>.+?)\s*"
r"(?:\[[^\]]+\])?\s*$"
)
def _severity(code: str, severity_word: str) -> str:
if code.startswith("SCS"):
return "high"
if severity_word == "error":
return "medium"
return "low"
def parse_dotnet_build_output(stdout: str, repo_root: str) -> list[Finding]:
out: list[Finding] = []
seen: set[tuple[str, int, str]] = set()
for raw in stdout.splitlines():
m = _LINE_RE.match(raw)
if not m:
continue
file_abs = m.group("file")
rel = os.path.relpath(file_abs, repo_root) if file_abs.startswith(repo_root) else file_abs
code = m.group("code")
line = int(m.group("line"))
key = (rel, line, code)
if key in seen:
continue
seen.add(key)
out.append(Finding(
tool="dotnet",
rule_id=code,
severity=_severity(code, m.group("severity")),
file=rel,
line=line,
end_line=line,
message=m.group("message"),
))
return out
@@ -0,0 +1,40 @@
"""Adapter: eslint -f json normalized to Finding[]."""
from __future__ import annotations
import json
import os
from scripts.manifest import Finding
def _severity(rule_id: str, sev_num: int) -> str:
if rule_id and "security" in rule_id:
return "high"
if sev_num == 2:
return "medium"
return "low"
def parse_eslint_output(stdout: str, repo_root: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for entry in payload:
file_path = entry.get("filePath", "")
rel = os.path.relpath(file_path, repo_root) if file_path.startswith(repo_root) else file_path
for msg in entry.get("messages", []):
rule_id = msg.get("ruleId") or ""
out.append(Finding(
tool="eslint",
rule_id=rule_id,
severity=_severity(rule_id, msg.get("severity", 1)),
file=rel,
line=msg.get("line", 0),
end_line=msg.get("endLine", msg.get("line", 0)),
message=msg.get("message", ""),
))
return out
@@ -0,0 +1,27 @@
"""Adapter: gitleaks --report-format json normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_gitleaks_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for r in payload:
out.append(Finding(
tool="gitleaks",
rule_id=r.get("RuleID", ""),
severity="critical",
file=r.get("File", ""),
line=r.get("StartLine", 0),
end_line=r.get("EndLine", r.get("StartLine", 0)),
message=r.get("Description", ""),
))
return out
@@ -0,0 +1,41 @@
"""Adapter: interrogate --quiet --output-format=json output → Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_interrogate_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, dict):
return []
files = payload.get("files")
if not isinstance(files, dict):
return []
out: list[Finding] = []
for file_path, file_data in files.items():
if not isinstance(file_data, dict):
continue
for entry in file_data.get("missing", []):
if entry.get("private"):
continue
full_name = entry.get("name", "")
symbol = full_name.split(":", 1)[-1] if ":" in full_name else full_name
if symbol.startswith("_"):
continue
kind = entry.get("type", "symbol")
line = entry.get("lineno", 0)
out.append(Finding(
tool="interrogate",
rule_id="interrogate:missing-docstring",
severity="low",
file=file_path,
line=line,
end_line=line,
message=f"missing docstring for public {kind} {symbol}",
))
return out
@@ -0,0 +1,34 @@
"""Adapter: jscpd JSON output → Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_jscpd_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, dict):
return []
out: list[Finding] = []
for dup in payload.get("duplicates", []):
first = dup.get("firstFile", {})
second = dup.get("secondFile", {})
lines = dup.get("lines", 0)
severity = "medium" if lines >= 30 else "low"
out.append(Finding(
tool="jscpd",
rule_id=f"jscpd:clone-{lines}lines",
severity=severity,
file=first.get("name", ""),
line=first.get("start", 0),
end_line=first.get("end", 0),
message=(
f"duplicate of {second.get('name', '')}:"
f"{second.get('start', 0)}-{second.get('end', 0)} ({lines} lines)"
),
))
return out
@@ -0,0 +1,31 @@
"""Adapter: knip --reporter json output → Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_knip_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, dict):
return []
out: list[Finding] = []
for issue in payload.get("issues", []):
file_path = issue.get("file", "")
for export in issue.get("exports", []):
name = export.get("name", "")
line = export.get("line", 0)
out.append(Finding(
tool="knip",
rule_id="knip:dead-export",
severity="medium",
file=file_path,
line=line,
end_line=line,
message=f"unused export '{name}'",
))
return out
@@ -0,0 +1,46 @@
"""Adapter: lizard --csv output → Finding[]."""
from __future__ import annotations
from scripts.manifest import Finding
def _severity_for_ccn(ccn: int) -> str | None:
if ccn >= 20:
return "high"
if ccn >= 10:
return "medium"
return None
def parse_lizard_output(stdout: str) -> list[Finding]:
out: list[Finding] = []
for raw in stdout.splitlines():
parts = raw.strip().split(",")
if len(parts) < 6:
continue
try:
ccn = int(parts[1])
except ValueError:
continue
severity = _severity_for_ccn(ccn)
if severity is None:
continue
location = parts[5]
loc_parts = location.split("@")
if len(loc_parts) < 3:
continue
name, line_str, file_path = loc_parts[0], loc_parts[1], loc_parts[2]
try:
line = int(line_str)
except ValueError:
continue
out.append(Finding(
tool="lizard",
rule_id=f"lizard:ccn={ccn}",
severity=severity,
file=file_path,
line=line,
end_line=line,
message=f"function {name} has CCN {ccn}",
))
return out
@@ -0,0 +1,32 @@
"""Adapter: luac -p stderr normalized to Finding[]."""
from __future__ import annotations
import re
from scripts.manifest import Finding
_PATTERN = re.compile(r"^luac:\s+(.+?):(\d+):\s+(.+)$")
def parse_luac_output(stderr: str, repo_root: str = "") -> list[Finding]:
out: list[Finding] = []
prefix = repo_root.rstrip("/") + "/" if repo_root else ""
for line in stderr.splitlines():
m = _PATTERN.match(line.strip())
if not m:
continue
filepath, lineno_str, message = m.group(1), m.group(2), m.group(3)
if prefix and filepath.startswith(prefix):
filepath = filepath[len(prefix):]
lineno = int(lineno_str)
out.append(Finding(
tool="luac",
rule_id="syntax-error",
severity="critical",
file=filepath,
line=lineno,
end_line=lineno,
message=message,
))
return out
@@ -0,0 +1,37 @@
"""Adapter: mypy text output normalized to Finding[]."""
from __future__ import annotations
import re
from scripts.manifest import Finding
_LINE_RE = re.compile(
r"^(?P<file>[^:]+):(?P<line>\d+):\s*"
r"(?P<severity>error|warning|note):\s*"
r"(?P<message>.+?)"
r"(?:\s+\[(?P<code>[a-z\-]+)\])?\s*$"
)
_SEVERITY = {"error": "medium", "warning": "low", "note": None}
def parse_mypy_output(stdout: str) -> list[Finding]:
out: list[Finding] = []
for raw in stdout.splitlines():
m = _LINE_RE.match(raw)
if not m:
continue
sev = _SEVERITY.get(m.group("severity"))
if sev is None:
continue
out.append(Finding(
tool="mypy",
rule_id=m.group("code") or "",
severity=sev,
file=m.group("file"),
line=int(m.group("line")),
end_line=int(m.group("line")),
message=m.group("message"),
))
return out
@@ -0,0 +1,51 @@
"""Adapter: opengrep --json normalized to Finding[]. Same JSON schema as semgrep (opengrep is a semgrep fork)."""
from __future__ import annotations
import json
import re
from scripts.manifest import Finding
_SEVERITY = {"ERROR": "high", "WARNING": "medium", "INFO": "low"}
_CWE_RE = re.compile(r"(CWE-\d+)")
_LOCAL_RULES_MARKER = "scripts.rules."
def _clean_rule_id(check_id: str) -> str:
if _LOCAL_RULES_MARKER in check_id:
return check_id.rsplit(_LOCAL_RULES_MARKER, 1)[-1]
return check_id
def _cwe(meta: dict) -> str | None:
raw = meta.get("cwe")
if isinstance(raw, list) and raw:
m = _CWE_RE.search(str(raw[0]))
return m.group(1) if m else None
if isinstance(raw, str):
m = _CWE_RE.search(raw)
return m.group(1) if m else None
return None
def parse_opengrep_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
out: list[Finding] = []
for r in payload.get("results", []):
extra = r.get("extra", {})
metadata = extra.get("metadata", {})
out.append(Finding(
tool="opengrep",
rule_id=_clean_rule_id(r.get("check_id", "")),
severity=_SEVERITY.get(extra.get("severity", ""), "medium"),
file=r.get("path", ""),
line=r.get("start", {}).get("line", 0),
end_line=r.get("end", {}).get("line", r.get("start", {}).get("line", 0)),
message=extra.get("message", ""),
cwe=_cwe(metadata),
))
return out
@@ -0,0 +1,31 @@
"""Adapter: osv-scanner --format json normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_osv_scanner_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
out: list[Finding] = []
for r in payload.get("results", []):
manifest_path = r.get("source", {}).get("path", "")
for pkg in r.get("packages", []):
p = pkg.get("package", {})
name = p.get("name", "")
version = p.get("version", "")
for v in pkg.get("vulnerabilities", []):
out.append(Finding(
tool="osv-scanner",
rule_id=v.get("id", ""),
severity="high",
file=manifest_path,
line=1,
end_line=1,
message=f"{name} {version}: {v.get('summary', '')}",
))
return out
@@ -0,0 +1,35 @@
"""Adapter: pip-audit -f json normalized to Finding[].
Findings are attached to the dependency manifest file rather than a source
file, since the vulnerability is in a pinned dep, not in code.
"""
from __future__ import annotations
import json
from scripts.manifest import Finding
def parse_pip_audit_output(stdout: str, manifest_path: str = "requirements.txt") -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
out: list[Finding] = []
for dep in payload.get("dependencies", []):
name = dep.get("name", "")
version = dep.get("version", "")
for v in dep.get("vulns", []):
fix = v.get("fix_versions") or []
fix_str = ", ".join(fix) if fix else None
out.append(Finding(
tool="pip-audit",
rule_id=v.get("id", ""),
severity="high",
file=manifest_path,
line=1,
end_line=1,
message=f"{name} {version}: {v.get('description', '')}",
fix_suggestion=f"upgrade to {fix_str}" if fix_str else None,
))
return out
@@ -0,0 +1,56 @@
"""Adapter: Invoke-ScriptAnalyzer JSON normalized to Finding[].
Shared by both PowerShell tools — PSScriptAnalyzer's built-in rules and the
InjectionHunter custom rule pack — because both are Invoke-ScriptAnalyzer runs
and emit the same DiagnosticRecord shape.
"""
from __future__ import annotations
import json
import os
from scripts.manifest import Finding
_SEVERITY = {
"ParseError": "critical",
"Error": "high",
"Warning": "medium",
"Information": "low",
}
_INJECTION_FLOOR = "high"
def parse_psscriptanalyzer_output(
stdout: str, repo_root: str, tool: str = "psscriptanalyzer",
) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if isinstance(payload, dict):
payload = [payload]
if not isinstance(payload, list):
return []
out: list[Finding] = []
for item in payload:
if not isinstance(item, dict):
continue
path = item.get("ScriptPath") or item.get("ScriptName") or ""
rel = os.path.relpath(path, repo_root) if path.startswith(repo_root) else path
line = item.get("Line") or 0
severity = _SEVERITY.get(item.get("Severity", ""), "medium")
if tool == "injectionhunter":
severity = _INJECTION_FLOOR
out.append(Finding(
tool=tool,
rule_id=item.get("RuleName", "unknown"),
severity=severity,
file=rel,
line=line,
end_line=item.get("EndLine") or line,
message=item.get("Message", ""),
))
return out
@@ -0,0 +1,46 @@
"""Adapter: radon cc -j JSON output → Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def _severity_for_complexity(cc: int) -> str | None:
if cc >= 20:
return "high"
if cc >= 10:
return "medium"
return None
def parse_radon_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, dict):
return []
out: list[Finding] = []
for file_path, items in payload.items():
if not isinstance(items, list):
continue
for item in items:
cc = item.get("complexity", 0)
severity = _severity_for_complexity(cc)
if severity is None:
continue
line = item.get("lineno", 0)
end_line = item.get("endline", line)
name = item.get("name", "?")
kind = item.get("type", "function")
out.append(Finding(
tool="radon",
rule_id=f"radon:cc={cc}",
severity=severity,
file=file_path,
line=line,
end_line=end_line,
message=f"high cyclomatic complexity (CCN={cc}) in {kind} {name}",
))
return out
@@ -0,0 +1,40 @@
"""Adapter: ruff check --output-format=json normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def _severity_for_rule(code: str) -> str:
if not code:
return "low"
if code.startswith("S"):
return "high"
if code.startswith("B"):
return "medium"
return "low"
def parse_ruff_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for item in payload:
loc = item.get("location") or {}
end = item.get("end_location") or loc
code = item.get("code", "")
out.append(Finding(
tool="ruff",
rule_id=code,
severity=_severity_for_rule(code),
file=item.get("filename", ""),
line=loc.get("row", 0),
end_line=end.get("row", loc.get("row", 0)),
message=item.get("message", ""),
))
return out
@@ -0,0 +1,38 @@
"""Adapter: ruff check with idiom/simplification selectors → Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
def _severity_for_idiom_rule(code: str) -> str:
if not code:
return "low"
if code.startswith(("PLR", "C90", "B")):
return "medium"
return "low"
def parse_ruff_idiom_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for item in payload:
loc = item.get("location") or {}
end = item.get("end_location") or loc
code = item.get("code", "")
out.append(Finding(
tool="ruff-idiom",
rule_id=code,
severity=_severity_for_idiom_rule(code),
file=item.get("filename", ""),
line=loc.get("row", 0),
end_line=end.get("row", loc.get("row", 0)),
message=item.get("message", ""),
))
return out
@@ -0,0 +1,34 @@
"""Adapter: selene --display-style=json normalized to Finding[]."""
from __future__ import annotations
import json
from scripts.manifest import Finding
_SEVERITY = {"Error": "high", "Warning": "medium", "Note": "low"}
def parse_selene_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
if not isinstance(payload, list):
return []
out: list[Finding] = []
for item in payload:
code = item.get("code", {})
span = item.get("span", {})
start = span.get("start", {})
end_span = span.get("end", {})
out.append(Finding(
tool="selene",
rule_id=code.get("name", "unknown"),
severity=_SEVERITY.get(code.get("severity", ""), "medium"),
file=item.get("filename", ""),
line=start.get("line", 0),
end_line=end_span.get("line", start.get("line", 0)),
message=item.get("primary_label", ""),
))
return out
@@ -0,0 +1,33 @@
"""Adapter: tsc --noEmit text output normalized to Finding[]."""
from __future__ import annotations
import re
from scripts.manifest import Finding
_LINE_RE = re.compile(
r"^(?P<file>[^(]+)\((?P<line>\d+),\d+\):\s*"
r"(?P<severity>error|warning)\s+"
r"(?P<code>TS\d+):\s*"
r"(?P<message>.+?)\s*$"
)
def parse_tsc_output(stdout: str) -> list[Finding]:
out: list[Finding] = []
for raw in stdout.splitlines():
m = _LINE_RE.match(raw)
if not m:
continue
sev = "medium" if m.group("severity") == "error" else "low"
out.append(Finding(
tool="tsc",
rule_id=m.group("code"),
severity=sev,
file=m.group("file"),
line=int(m.group("line")),
end_line=int(m.group("line")),
message=m.group("message"),
))
return out
@@ -0,0 +1,31 @@
"""Adapter: vulture text output → Finding[]."""
from __future__ import annotations
import re
from scripts.manifest import Finding
_LINE_RE = re.compile(
r"^(?P<file>[^:]+):(?P<line>\d+):\s*(?P<msg>.+?)\s*\((?P<conf>\d+)%\s*confidence\)$"
)
def parse_vulture_output(stdout: str) -> list[Finding]:
out: list[Finding] = []
for raw in stdout.splitlines():
m = _LINE_RE.match(raw.strip())
if not m:
continue
conf = int(m.group("conf"))
severity = "medium" if conf >= 80 else "low"
out.append(Finding(
tool="vulture",
rule_id=f"vulture:{conf}pct",
severity=severity,
file=m.group("file"),
line=int(m.group("line")),
end_line=int(m.group("line")),
message=m.group("msg"),
))
return out
@@ -0,0 +1,48 @@
"""Adapter: zizmor --format sarif normalized to Finding[]."""
from __future__ import annotations
import json
import re
from scripts.manifest import Finding
_LEVEL = {"error": "high", "warning": "medium", "note": "low"}
_CWE_RE = re.compile(r"(CWE-\d+)")
def parse_zizmor_output(stdout: str) -> list[Finding]:
try:
payload = json.loads(stdout)
except json.JSONDecodeError:
return []
out: list[Finding] = []
for run in payload.get("runs", []):
cwe_by_rule: dict[str, str] = {}
for rule in run.get("tool", {}).get("driver", {}).get("rules", []):
for tag in rule.get("properties", {}).get("tags", []):
m = _CWE_RE.match(tag)
if m:
cwe_by_rule[rule["id"]] = m.group(1)
break
for result in run.get("results", []):
rule_id = result.get("ruleId", "unknown")
locations = result.get("locations", [])
if not locations:
continue
phys = locations[0].get("physicalLocation", {})
uri = phys.get("artifactLocation", {}).get("uri", "")
region = phys.get("region", {})
start_line = region.get("startLine", 0)
end_line = region.get("endLine", start_line)
out.append(Finding(
tool="zizmor",
rule_id=rule_id,
severity=_LEVEL.get(result.get("level", "warning"), "medium"),
file=uri,
line=start_line,
end_line=end_line,
message=result.get("message", {}).get("text", ""),
cwe=cwe_by_rule.get(rule_id),
))
return out
@@ -0,0 +1,464 @@
"""collect-findings.py — produce a audit-code manifest.json."""
from __future__ import annotations
import argparse
import glob
import json
import subprocess
import sys
from collections import Counter
from pathlib import Path
_HERE = Path(__file__).resolve().parent
if str(_HERE.parent) not in sys.path:
sys.path.insert(0, str(_HERE.parent))
from scripts.adapters import (
bandit as ad_bandit,
ruff as ad_ruff,
ruff_idiom as ad_ruff_idiom,
mypy as ad_mypy,
eslint as ad_eslint,
tsc as ad_tsc,
opengrep as ad_opengrep,
gitleaks as ad_gitleaks,
pip_audit as ad_pip_audit,
osv_scanner as ad_osv,
dotnet as ad_dotnet,
vulture as ad_vulture,
radon as ad_radon,
interrogate as ad_interrogate,
lizard as ad_lizard,
knip as ad_knip,
jscpd as ad_jscpd,
selene as ad_selene,
luac as ad_luac,
psscriptanalyzer as ad_pssa,
actionlint as ad_actionlint,
zizmor as ad_zizmor,
)
from scripts.diff_filter import filter_findings_by_diff
from scripts.git_diff import (
changed_files_with_ranges, resolve_base_ref, resolve_default_branch,
)
from scripts.language_detect import (
detect_language, is_dep_manifest, is_gha_file, is_supported_source,
known_unsupported_language,
)
from scripts.manifest import (
ChangedFile, Finding, LanguageBreakdown, Manifest,
PackageDiff, ToolStat,
)
from scripts.package_diff import diff_package_json, diff_requirements_txt
from scripts.runner import run_tool, tool_available
from scripts.slicing import slice_for_agent
def _git_show(repo: str, ref: str, path: str) -> str:
"""Return file content at `ref`, or empty string if not present (e.g. new file)."""
r = subprocess.run(
["git", "-C", repo, "show", f"{ref}:{path}"],
capture_output=True, text=True, check=False,
)
return r.stdout if r.returncode == 0 else ""
def _build_package_diffs(
repo: str, base: str, dep_manifest_paths: list[str],
) -> dict[str, PackageDiff]:
diffs: dict[str, PackageDiff] = {
"python": PackageDiff(),
"javascript": PackageDiff(),
"csharp": PackageDiff(),
}
for path in dep_manifest_paths:
name = Path(path).name
before = _git_show(repo, base, path)
after_path = Path(repo) / path
after = after_path.read_text() if after_path.is_file() else ""
if name.endswith(".txt") and "requirements" in name:
pd = diff_requirements_txt(before, after)
_merge_package_diff(diffs["python"], pd)
elif name == "package.json":
pd = diff_package_json(before, after)
_merge_package_diff(diffs["javascript"], pd)
return diffs
def _merge_package_diff(into: PackageDiff, src: PackageDiff) -> None:
into.added.extend(src.added)
into.removed.extend(src.removed)
into.upgraded.extend(src.upgraded)
def _build_breakdown(changed_files: list[ChangedFile], skipped: list[str]) -> LanguageBreakdown:
n = Counter(cf.language for cf in changed_files)
return LanguageBreakdown(
python=n["python"], javascript=n["javascript"], typescript=n["typescript"],
csharp=n["csharp"], lua=n["lua"], powershell=n["powershell"],
github_actions=n["github-actions"], skipped_files=skipped,
)
def _run_python_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
for binary, args, adapter in [
("bandit", ["bandit", "-r", ".", "-f", "json", "-q"], ad_bandit.parse_bandit_output),
("ruff", ["ruff", "check", "--output-format=json", "."], ad_ruff.parse_ruff_output),
("ruff-idiom", [
"ruff", "check",
"--select", "SIM,PERF,UP,RET,PLR,C90,B",
"--output-format=json",
"--isolated",
".",
], ad_ruff_idiom.parse_ruff_idiom_output),
("mypy", ["mypy", "."], ad_mypy.parse_mypy_output),
]:
if not tool_available(binary):
stats[binary] = ToolStat(ran=False, reason="not on PATH")
continue
r = run_tool(args, cwd=repo)
parsed = adapter(r.stdout)
stats[binary] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
return findings, stats
def _run_js_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("eslint"):
r = run_tool(["eslint", ".", "-f", "json"], cwd=repo)
parsed = ad_eslint.parse_eslint_output(r.stdout, repo_root=repo)
stats["eslint"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["eslint"] = ToolStat(ran=False, reason="not on PATH")
if tool_available("tsc"):
r = run_tool(["tsc", "--noEmit"], cwd=repo)
parsed = ad_tsc.parse_tsc_output(r.stdout)
stats["tsc"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["tsc"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
def _run_dotnet_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("dotnet"):
r = run_tool(["dotnet", "build", "--no-incremental"], cwd=repo)
parsed = ad_dotnet.parse_dotnet_build_output(r.stdout, repo_root=repo)
stats["dotnet"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["dotnet"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
def _run_lua_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("selene"):
r = run_tool(["selene", "--display-style=json", "."], cwd=repo)
parsed = ad_selene.parse_selene_output(r.stdout)
stats["selene"] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
else:
stats["selene"] = ToolStat(ran=False, reason="not on PATH")
if tool_available("luac"):
lua_files = sorted(glob.glob(f"{repo}/**/*.lua", recursive=True))
combined_stderr = ""
for lua_file in lua_files:
r = run_tool(["luac", "-p", lua_file], cwd=repo)
if r.stderr:
combined_stderr += r.stderr
parsed = ad_luac.parse_luac_output(combined_stderr, repo_root=repo)
stats["luac"] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
else:
stats["luac"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
_PSSA_PROJECT = (
" | Select-Object RuleName,"
"@{n='Severity';e={$_.Severity.ToString()}},"
"ScriptPath,"
"@{n='Line';e={$_.Extent.StartLineNumber}},"
"@{n='EndLine';e={$_.Extent.EndLineNumber}},"
"Message | ConvertTo-Json -Depth 3 -AsArray"
)
_PSSA_COMMAND = "Invoke-ScriptAnalyzer -Path . -Recurse -ErrorAction SilentlyContinue"
_INJECTION_HUNTER_COMMAND = (
"$m = (Get-Module -ListAvailable -Name InjectionHunter | Select-Object -First 1).Path; "
"Invoke-ScriptAnalyzer -Path . -Recurse -CustomRulePath $m -ErrorAction SilentlyContinue"
)
def _pwsh_module_available(repo: str, module: str) -> bool:
r = run_tool(
["pwsh", "-NoProfile", "-NonInteractive", "-Command",
f"if (Get-Module -ListAvailable -Name {module}) {{ exit 0 }} else {{ exit 1 }}"],
cwd=repo,
)
return r.ran and r.exit_code == 0
def _run_powershell_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if not tool_available("pwsh"):
for name in ("psscriptanalyzer", "injectionhunter"):
stats[name] = ToolStat(ran=False, reason="pwsh not on PATH")
return findings, stats
for tool, module, command in [
("psscriptanalyzer", "PSScriptAnalyzer", _PSSA_COMMAND),
("injectionhunter", "InjectionHunter", _INJECTION_HUNTER_COMMAND),
]:
if not _pwsh_module_available(repo, module):
stats[tool] = ToolStat(ran=False, reason=f"{module} module not installed")
continue
r = run_tool(
["pwsh", "-NoProfile", "-NonInteractive", "-Command", command + _PSSA_PROJECT],
cwd=repo,
)
parsed = ad_pssa.parse_psscriptanalyzer_output(r.stdout, repo_root=repo, tool=tool)
stats[tool] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
return findings, stats
def _run_gha_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("actionlint"):
r = run_tool(["actionlint", "-format", "{{json .}}", ".github/"], cwd=repo)
parsed = ad_actionlint.parse_actionlint_output(r.stdout)
stats["actionlint"] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
else:
stats["actionlint"] = ToolStat(ran=False, reason="not on PATH")
if tool_available("zizmor"):
r = run_tool(["zizmor", "--format", "sarif", ".github/"], cwd=repo)
parsed = ad_zizmor.parse_zizmor_output(r.stdout)
stats["zizmor"] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
else:
stats["zizmor"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
def _run_polyglot_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("opengrep"):
lua_rules = _HERE / "rules" / "lua-security.yaml"
r = run_tool(
["opengrep", "--config=auto", f"--config={lua_rules}", "--json", "--quiet"],
cwd=repo,
)
parsed = ad_opengrep.parse_opengrep_output(r.stdout)
stats["opengrep"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["opengrep"] = ToolStat(ran=False, reason="not on PATH")
if tool_available("gitleaks"):
r = run_tool(["gitleaks", "detect", "--report-format=json", "--report-path=/dev/stdout", "--no-banner"], cwd=repo)
parsed = ad_gitleaks.parse_gitleaks_output(r.stdout)
stats["gitleaks"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["gitleaks"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
def _run_maintainability_tools(
repo: str, languages_present: set[str]
) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
is_python = "python" in languages_present
is_js = "javascript" in languages_present or "typescript" in languages_present
for binary, args, adapter, should_run in [
("vulture", ["vulture", "."], ad_vulture.parse_vulture_output, is_python),
("radon", ["radon", "cc", "-j", "."], ad_radon.parse_radon_output, is_python),
("interrogate", ["interrogate", "--quiet", "--output-format=json", "."], ad_interrogate.parse_interrogate_output, is_python),
("lizard", ["lizard", "--csv", "."], ad_lizard.parse_lizard_output, True),
("npx", ["npx", "knip", "--reporter", "json"], ad_knip.parse_knip_output, is_js),
("jscpd", ["jscpd", "--reporters", "json", "--silent", "."], ad_jscpd.parse_jscpd_output, is_js),
]:
tool_name = "knip" if binary == "npx" else binary
if not should_run:
continue
if not tool_available(binary):
stats[tool_name] = ToolStat(ran=False, reason="not on PATH")
continue
r = run_tool(args, cwd=repo)
parsed = adapter(r.stdout)
stats[tool_name] = ToolStat(ran=True, pre_filter=len(parsed), post_filter=0)
findings.extend(parsed)
return findings, stats
def _run_dep_tools(repo: str) -> tuple[list[Finding], dict[str, ToolStat]]:
findings: list[Finding] = []
stats: dict[str, ToolStat] = {}
if tool_available("pip-audit"):
r = run_tool(["pip-audit", "-f", "json"], cwd=repo)
parsed = ad_pip_audit.parse_pip_audit_output(r.stdout)
stats["pip-audit"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["pip-audit"] = ToolStat(ran=False, reason="not on PATH")
if tool_available("osv-scanner"):
r = run_tool(["osv-scanner", "--recursive", "--format=json", "."], cwd=repo)
parsed = ad_osv.parse_osv_scanner_output(r.stdout)
stats["osv-scanner"] = ToolStat(ran=True, pre_filter=len(parsed))
findings.extend(parsed)
else:
stats["osv-scanner"] = ToolStat(ran=False, reason="not on PATH")
return findings, stats
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--repo", required=True)
parser.add_argument("--base", default=None)
parser.add_argument("--head", default="HEAD")
parser.add_argument("--output-dir", required=True)
parser.add_argument("--mode", choices=("local", "ref"), required=True)
args = parser.parse_args(argv)
repo = str(Path(args.repo).resolve())
out_dir = Path(args.output_dir).resolve()
out_dir.mkdir(parents=True, exist_ok=True)
default_branch = resolve_default_branch(repo)
base = args.base or resolve_base_ref(repo, default_branch)
errors: list[str] = []
try:
changed = changed_files_with_ranges(repo, base, args.head)
except RuntimeError as e:
manifest = Manifest(
mode=args.mode, base_ref=base, head_ref=args.head,
default_branch=default_branch,
language_breakdown=LanguageBreakdown(),
changed_files=[], findings=[],
package_diffs={}, tool_stats={}, tools_unavailable=[],
errors=[str(e)],
)
(out_dir / "manifest.json").write_text(manifest.to_json())
return 1
changed_files: list[ChangedFile] = []
skipped: list[str] = []
dep_manifest_paths: list[str] = []
unsupported_by_language: dict[str, list[str]] = {}
for path, ranges in changed:
if is_supported_source(path):
changed_files.append(ChangedFile(
path=path,
language=detect_language(path),
added_lines=ranges,
))
continue
if is_gha_file(path):
changed_files.append(ChangedFile(
path=path,
language="github-actions",
added_lines=ranges,
))
continue
if is_dep_manifest(path):
dep_manifest_paths.append(path)
continue
unsupported = known_unsupported_language(path)
if unsupported is not None:
unsupported_by_language.setdefault(unsupported, []).append(path)
skipped.append(path)
for language, paths in sorted(unsupported_by_language.items()):
errors.append(
f"unsupported language not reviewed: {language} ({', '.join(paths)})"
)
breakdown = _build_breakdown(changed_files, skipped)
if not changed_files and not dep_manifest_paths:
errors.append("no supported source files in change set; skipped: " + ", ".join(skipped))
manifest = Manifest(
mode=args.mode, base_ref=base, head_ref=args.head,
default_branch=default_branch, language_breakdown=breakdown,
changed_files=[], findings=[],
package_diffs={}, tool_stats={}, tools_unavailable=[],
errors=errors,
)
(out_dir / "manifest.json").write_text(manifest.to_json())
return 1
findings: list[Finding] = []
tool_stats: dict[str, ToolStat] = {}
def collect(runner, *args) -> None:
f, s = runner(*args)
findings.extend(f)
tool_stats.update(s)
languages_present = {cf.language for cf in changed_files}
if "python" in languages_present:
collect(_run_python_tools, repo)
if "javascript" in languages_present or "typescript" in languages_present:
collect(_run_js_tools, repo)
if "csharp" in languages_present:
collect(_run_dotnet_tools, repo)
collect(_run_polyglot_tools, repo)
collect(_run_dep_tools, repo)
collect(_run_maintainability_tools, repo, languages_present)
if "lua" in languages_present:
collect(_run_lua_tools, repo)
if "powershell" in languages_present:
collect(_run_powershell_tools, repo)
if "github-actions" in languages_present:
collect(_run_gha_tools, repo)
filtered = filter_findings_by_diff(findings, changed_files, repo_root=repo)
for tool_name, stat in tool_stats.items():
if stat.ran:
stat.post_filter = sum(1 for f in filtered if f.tool == tool_name)
tools_unavailable = [name for name, s in tool_stats.items() if not s.ran]
package_diffs = _build_package_diffs(repo, base, dep_manifest_paths)
manifest = Manifest(
mode=args.mode,
base_ref=base,
head_ref=args.head,
default_branch=default_branch,
language_breakdown=breakdown,
changed_files=changed_files,
findings=filtered,
package_diffs=package_diffs,
tool_stats=tool_stats,
tools_unavailable=tools_unavailable,
errors=errors,
)
(out_dir / "manifest.json").write_text(manifest.to_json())
manifest_dict = manifest.to_dict()
for agent in ("security-triage", "type-safety", "dependency", "consistency", "secrets", "maintainability", "walkthrough", "gha-reviewer"):
sliced = slice_for_agent(manifest_dict, agent)
(out_dir / f"manifest-{agent}.json").write_text(json.dumps(sliced, indent=2))
return 0
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,41 @@
"""Filter Finding[] to those overlapping any changed-line range."""
from __future__ import annotations
import os
from scripts.manifest import ChangedFile, Finding
def _overlaps(a: tuple[int, int], b: tuple[int, int]) -> bool:
return not (a[1] < b[0] or b[1] < a[0])
def _normalize(path: str, repo_root: str | None = None) -> str:
"""Strip `./` prefix and (optionally) the repo root so adapter paths
line up with git-diff paths regardless of which form the tool emitted.
"""
if repo_root and path.startswith(repo_root):
path = os.path.relpath(path, repo_root)
while path.startswith("./"):
path = path[2:]
return path
def filter_findings_by_diff(
findings: list[Finding], changed_files: list[ChangedFile],
repo_root: str | None = None,
) -> list[Finding]:
by_path: dict[str, list[tuple[int, int]]] = {
_normalize(cf.path, repo_root): list(cf.added_lines)
for cf in changed_files
}
out: list[Finding] = []
for f in findings:
norm = _normalize(f.file, repo_root)
ranges = by_path.get(norm)
if not ranges:
continue
if any(_overlaps((f.line, f.end_line), r) for r in ranges):
f.file = norm
out.append(f)
return out
@@ -0,0 +1,122 @@
"""Git diff scanning with origin-base resolution and defensive flags."""
from __future__ import annotations
import re
import subprocess
_HUNK_RE = re.compile(r"^@@ -\d+(?:,\d+)? \+(\d+)(?:,(\d+))? @@")
def resolve_default_branch(repo: str) -> str:
try:
r = subprocess.run(
["git", "-C", repo, "symbolic-ref", "refs/remotes/origin/HEAD"],
capture_output=True, text=True, check=False,
)
if r.returncode == 0:
return r.stdout.strip().rsplit("/", 1)[-1]
except FileNotFoundError:
pass
for candidate in ("main", "master"):
r = subprocess.run(
["git", "-C", repo, "rev-parse", f"origin/{candidate}"],
capture_output=True, text=True, check=False,
)
if r.returncode == 0:
return candidate
return "main"
def resolve_base_ref(repo: str, branch: str) -> str:
"""Return the diff base for `branch`, fetching origin first."""
try:
subprocess.run(
["git", "-C", repo, "fetch", "--quiet", "--no-tags",
"origin", branch],
capture_output=True, text=True, check=False, timeout=60,
)
except (FileNotFoundError, subprocess.TimeoutExpired):
pass
check = subprocess.run(
["git", "-C", repo, "rev-parse", "--verify", "--quiet",
f"origin/{branch}"],
capture_output=True, text=True, check=False,
)
return f"origin/{branch}" if check.returncode == 0 else branch
def _run_git_diff(repo: str, base: str, head: str) -> str:
result = subprocess.run(
[
"git", "-C", repo,
"-c", "diff.noprefix=false",
"-c", "color.ui=never",
"diff", "--no-ext-diff", "--unified=0", f"{base}...{head}",
],
capture_output=True, text=True, check=False,
)
if result.returncode != 0:
raise RuntimeError(f"git diff failed: {result.stderr.strip()}")
return result.stdout
def _iter_file_blocks(diff_text: str):
current: str | None = None
buf: list[str] = []
for line in diff_text.splitlines():
if line.startswith("diff --git "):
if current is not None:
yield current, "\n".join(buf)
current = None
buf = []
elif line.startswith("+++ b/"):
current = line[len("+++ b/"):]
if current is not None:
buf.append(line)
if current is not None:
yield current, "\n".join(buf)
def _added_ranges(block: str) -> list[tuple[int, int]]:
ranges: list[tuple[int, int]] = []
cur: int | None = None
start: int | None = None
end: int | None = None
for line in block.splitlines():
m = _HUNK_RE.match(line)
if m:
if start is not None:
ranges.append((start, end)) # type: ignore[arg-type]
start = end = None
cur = int(m.group(1))
continue
if cur is None:
continue
if line.startswith("+") and not line.startswith("+++"):
if start is None:
start = cur
end = cur
cur += 1
elif line.startswith("-") and not line.startswith("---"):
continue
else:
if start is not None:
ranges.append((start, end)) # type: ignore[arg-type]
start = end = None
cur += 1
if start is not None:
ranges.append((start, end)) # type: ignore[arg-type]
return ranges
def changed_files_with_ranges(
repo: str, base: str, head: str,
) -> list[tuple[str, list[tuple[int, int]]]]:
"""Return [(path, added_line_ranges), ...] for every changed file."""
out: list[tuple[str, list[tuple[int, int]]]] = []
for path, block in _iter_file_blocks(_run_git_diff(repo, base, head)):
ranges = _added_ranges(block)
if ranges:
out.append((path, ranges))
return out
+160
View File
@@ -0,0 +1,160 @@
#!/usr/bin/env bash
set -euo pipefail
SKILL_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$SKILL_DIR"
say() { printf '\n\033[1m▶ %s\033[0m\n' "$*"; }
warn() { printf '\033[33m! %s\033[0m\n' "$*"; }
if ! command -v uv >/dev/null 2>&1; then
warn "uv not installed. Install: https://docs.astral.sh/uv/getting-started/installation/"
warn "Falling back to plain pip — tools will install into the active environment."
if ! command -v pip >/dev/null 2>&1; then
echo "Neither uv nor pip available. Aborting Python-tools install."
exit 1
fi
pip install bandit ruff mypy pip-audit
else
say "Installing Python tools into $SKILL_DIR/.venv/ via uv"
uv sync --group tools
fi
install_opengrep() {
if [[ -x "$SKILL_DIR/.venv/bin/opengrep" ]]; then
echo " opengrep: already installed (.venv/bin)"
return
fi
say "Installing opengrep into $SKILL_DIR/.venv/bin/"
local os arch asset
os="$(uname -s)"
arch="$(uname -m)"
case "$os-$arch" in
Linux-x86_64) asset="opengrep_manylinux_x86" ;;
Linux-aarch64) asset="opengrep_manylinux_aarch64" ;;
Darwin-x86_64) asset="opengrep_osx_x86" ;;
Darwin-arm64) asset="opengrep_osx_arm64" ;;
*)
warn "opengrep: unsupported platform $os-$arch, install manually from https://github.com/opengrep/opengrep/releases"
return
;;
esac
local tag url
tag="$(curl -fsSL https://api.github.com/repos/opengrep/opengrep/releases/latest | grep -m1 '"tag_name"' | sed -E 's/.*"([^"]+)".*/\1/')"
if [[ -z "$tag" ]]; then
warn "opengrep: could not resolve latest release tag, install manually"
return
fi
url="https://github.com/opengrep/opengrep/releases/download/$tag/$asset"
mkdir -p "$SKILL_DIR/.venv/bin"
curl -fsSL "$url" -o "$SKILL_DIR/.venv/bin/opengrep"
chmod +x "$SKILL_DIR/.venv/bin/opengrep"
}
install_opengrep
install_powershell_modules() {
if ! command -v pwsh >/dev/null 2>&1; then
warn "pwsh not found — PowerShell review (PSScriptAnalyzer, InjectionHunter) will be skipped."
warn " Install: https://learn.microsoft.com/powershell/scripting/install/installing-powershell"
return
fi
say "Installing PowerShell modules for the current user"
pwsh -NoProfile -NonInteractive -Command '
foreach ($m in "PSScriptAnalyzer", "InjectionHunter") {
if (Get-Module -ListAvailable -Name $m) {
Write-Host " $m: already installed"
} else {
Install-Module -Name $m -Scope CurrentUser -Force -AcceptLicense -Repository PSGallery
Write-Host " $m: installed"
}
}'
}
install_powershell_modules
install_native_brew() {
say "Installing native tools via Homebrew"
for pkg in gitleaks osv-scanner gh; do
if brew list --formula | grep -qx "$pkg"; then
echo " $pkg: already installed"
else
brew install "$pkg"
fi
done
}
install_native_apt() {
say "Installing native tools via apt-get"
if ! command -v gh >/dev/null 2>&1; then
warn "gh: follow https://github.com/cli/cli/blob/trunk/docs/install_linux.md"
fi
if ! command -v gitleaks >/dev/null 2>&1; then
warn "gitleaks: download from https://github.com/gitleaks/gitleaks/releases"
fi
if ! command -v osv-scanner >/dev/null 2>&1; then
warn "osv-scanner: install via 'go install github.com/google/osv-scanner/cmd/osv-scanner@latest' or download from https://github.com/google/osv-scanner/releases"
fi
}
install_native_arch() {
local helper
if command -v paru >/dev/null 2>&1; then
helper="paru"
elif command -v yay >/dev/null 2>&1; then
helper="yay"
else
helper="pacman"
fi
say "Installing native tools via $helper"
local pkgs=(gitleaks github-cli osv-scanner)
if [[ "$helper" == "pacman" ]]; then
sudo pacman -S --needed --noconfirm gitleaks github-cli || true
if ! command -v osv-scanner >/dev/null 2>&1; then
warn "osv-scanner is AUR-only; pacman can't install it. Use paru/yay or install manually:"
warn " go install github.com/google/osv-scanner/cmd/osv-scanner@latest"
fi
else
"$helper" -S --needed --noconfirm "${pkgs[@]}"
fi
}
if command -v brew >/dev/null 2>&1; then
install_native_brew
elif command -v pacman >/dev/null 2>&1; then
install_native_arch
elif command -v apt-get >/dev/null 2>&1; then
install_native_apt
else
warn "No supported native package manager found (brew/pacman/apt). Install gitleaks, osv-scanner, gh manually."
fi
say "Verifying tool availability"
for tool in bandit ruff mypy pip-audit opengrep vulture radon interrogate lizard gitleaks osv-scanner gh pwsh; do
if [[ -x "$SKILL_DIR/.venv/bin/$tool" ]]; then
printf ' %-15s %s\n' "$tool" "(.venv/bin)"
elif command -v "$tool" >/dev/null 2>&1; then
printf ' %-15s %s\n' "$tool" "$(command -v "$tool")"
else
printf ' %-15s \033[31mmissing\033[0m\n' "$tool"
fi
done
cat <<EOF
Per-project tools (not installed here — must live in the target repo):
- eslint + eslint-plugin-security (npm i -D)
- typescript (tsc) (npm i -D)
- SecurityCodeScan + dotnet (dotnet add package SecurityCodeScan.VS2019)
- knip (npm i -D)
- jscpd (npm i -D)
Run /audit-code to use the skill.
EOF
@@ -0,0 +1,111 @@
"""Detect the language of a changed file from its path."""
from __future__ import annotations
from pathlib import PurePosixPath
SOURCE_EXTENSIONS: dict[str, str] = {
".py": "python",
".js": "javascript",
".jsx": "javascript",
".mjs": "javascript",
".cjs": "javascript",
".ts": "typescript",
".tsx": "typescript",
".cs": "csharp",
".lua": "lua",
".ps1": "powershell",
".psm1": "powershell",
".psd1": "powershell",
}
DEP_MANIFESTS: set[str] = {
"requirements.txt", "pyproject.toml", "poetry.lock", "Pipfile.lock",
"Pipfile", "setup.py", "setup.cfg",
"package.json", "package-lock.json", "yarn.lock", "pnpm-lock.yaml",
"packages.config",
}
DEP_MANIFEST_PREFIXES: tuple[str, ...] = ("requirements",)
DEP_MANIFEST_EXTENSIONS: tuple[str, ...] = (".csproj", ".sln")
KNOWN_UNSUPPORTED_LANGUAGES: dict[str, str] = {
".go": "Go",
".rb": "Ruby",
".java": "Java",
".kt": "Kotlin",
".kts": "Kotlin",
".rs": "Rust",
".swift": "Swift",
".php": "PHP",
".scala": "Scala",
".sc": "Scala",
".clj": "Clojure",
".cljs": "ClojureScript",
".ex": "Elixir",
".exs": "Elixir",
".erl": "Erlang",
".hrl": "Erlang",
".dart": "Dart",
".r": "R",
".jl": "Julia",
".zig": "Zig",
".nim": "Nim",
".hs": "Haskell",
".ml": "OCaml",
".mli": "OCaml",
".fs": "F#",
".fsi": "F#",
".fsx": "F#",
".vb": "Visual Basic",
".pl": "Perl",
".pm": "Perl",
".c": "C",
".h": "C/C++ header",
".cpp": "C++",
".cc": "C++",
".cxx": "C++",
".hpp": "C++",
".m": "Objective-C",
".mm": "Objective-C++",
".groovy": "Groovy",
}
_GHA_YAML_EXTS: frozenset[str] = frozenset((".yml", ".yaml"))
_GHA_PATH_PREFIXES: tuple[str, ...] = (".github/workflows/", ".github/actions/")
def detect_language(path: str) -> str | None:
suffix = PurePosixPath(path).suffix
return SOURCE_EXTENSIONS.get(suffix)
def is_supported_source(path: str) -> bool:
return detect_language(path) is not None
def is_gha_file(path: str) -> bool:
"""Return True if path is a GitHub Actions workflow or action definition."""
normalized = path.replace("\\", "/")
if PurePosixPath(normalized).suffix.lower() not in _GHA_YAML_EXTS:
return False
return normalized.startswith(_GHA_PATH_PREFIXES)
def known_unsupported_language(path: str) -> str | None:
"""Return the display name of a recognized-but-unsupported source language,
or None if the path isn't one we recognize as source code.
"""
suffix = PurePosixPath(path).suffix.lower()
return KNOWN_UNSUPPORTED_LANGUAGES.get(suffix)
def is_dep_manifest(path: str) -> bool:
name = PurePosixPath(path).name
if name in DEP_MANIFESTS:
return True
if any(name.startswith(p) and name.endswith(".txt") for p in DEP_MANIFEST_PREFIXES):
return True
return any(name.endswith(ext) for ext in DEP_MANIFEST_EXTENSIONS)
@@ -0,0 +1,84 @@
"""Append one subagent_run row per agent after the fan-out completes.
The orchestrator calls this once after all subagents return. It scans
<output-dir>/findings-<agent>.json for each agent listed in the
--usage-json payload, counts findings, and writes a subagent_run row to
runs.jsonl using token / duration metadata supplied by the orchestrator.
Usage:
python scripts/log-run.py \\
--output-dir <OUTPUT> --run-id <hex> --repo <path> --mode <local|ref> \\
--usage-json - <<JSON
{
"walkthrough-reviewer": {"model":"sonnet","input_tokens":1234,"output_tokens":567,"duration_ms":4500},
"security-triage-reviewer":{"model":"sonnet","input_tokens":2345,"output_tokens":678,"duration_ms":5200}
}
JSON
`--log-path` defaults to ~/.claude/cache/audit-code/runs.jsonl.
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
from scripts.telemetry import append_subagent_run
_DEFAULT_LOG = Path.home() / ".claude/cache/audit-code/runs.jsonl"
def _count_findings(path: Path) -> int:
if not path.exists():
return 0
try:
data = json.loads(path.read_text(encoding="utf-8"))
except json.JSONDecodeError:
return 0
if isinstance(data, dict):
findings = data.get("findings")
if isinstance(findings, list):
return len(findings)
elif isinstance(data, list):
return len(data)
return 0
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--output-dir", required=True, type=Path)
p.add_argument("--run-id", required=True)
p.add_argument("--repo", required=True)
p.add_argument("--mode", required=True, choices=["local", "ref"])
p.add_argument("--log-path", type=Path, default=_DEFAULT_LOG)
p.add_argument("--usage-json", required=True,
help="Path to JSON file, or '-' for stdin.")
args = p.parse_args(argv)
raw = sys.stdin.read() if args.usage_json == "-" else Path(args.usage_json).read_text(encoding="utf-8")
usage = json.loads(raw)
if not isinstance(usage, dict):
print("usage-json must be a JSON object keyed by agent name", file=sys.stderr)
return 2
for agent, meta in usage.items():
if not isinstance(meta, dict):
print(f"skipping {agent}: usage entry not an object", file=sys.stderr)
continue
findings_path = args.output_dir / f"findings-{agent}.json"
append_subagent_run(
args.log_path,
run_id=args.run_id, repo=args.repo, mode=args.mode, agent=agent,
model=str(meta.get("model", "?")),
input_tokens=int(meta.get("input_tokens", 0)),
output_tokens=int(meta.get("output_tokens", 0)),
duration_ms=int(meta.get("duration_ms", 0)),
finding_count=_count_findings(findings_path),
)
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
@@ -0,0 +1,146 @@
"""Manifest dataclasses for collect-findings.py output.
The manifest is the contract between the script and the review subagents.
Schema mirrors DESIGN.md.
"""
from __future__ import annotations
import json
from dataclasses import dataclass, field, asdict
from typing import Literal
Mode = Literal["local", "ref"]
Severity = Literal["critical", "high", "medium", "low", "info"]
Language = Literal[
"python", "javascript", "typescript", "csharp", "lua", "powershell",
"github-actions",
]
@dataclass
class Finding:
tool: str
rule_id: str
severity: Severity
file: str
line: int
end_line: int
message: str
cwe: str | None = None
fix_suggestion: str | None = None
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class ChangedFile:
path: str
language: str
added_lines: list[tuple[int, int]]
def to_dict(self) -> dict:
return {
"path": self.path,
"language": self.language,
"added_lines": [list(r) for r in self.added_lines],
}
@dataclass
class LanguageBreakdown:
python: int = 0
javascript: int = 0
typescript: int = 0
csharp: int = 0
lua: int = 0
powershell: int = 0
github_actions: int = 0
skipped_files: list[str] = field(default_factory=list)
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class PackageEntry:
name: str
version: str
def to_dict(self) -> dict:
return asdict(self)
@dataclass
class PackageUpgrade:
name: str
from_version: str
to_version: str
def to_dict(self) -> dict:
return {"name": self.name, "from": self.from_version, "to": self.to_version}
@dataclass
class PackageDiff:
added: list[PackageEntry] = field(default_factory=list)
removed: list[PackageEntry] = field(default_factory=list)
upgraded: list[PackageUpgrade] = field(default_factory=list)
def to_dict(self) -> dict:
return {
"added": [p.to_dict() for p in self.added],
"removed": [p.to_dict() for p in self.removed],
"upgraded": [u.to_dict() for u in self.upgraded],
}
@dataclass
class ToolStat:
ran: bool
pre_filter: int = 0
post_filter: int = 0
reason: str | None = None
def to_dict(self) -> dict:
d: dict = {"ran": self.ran}
if self.ran:
d["pre_filter"] = self.pre_filter
d["post_filter"] = self.post_filter
else:
d["reason"] = self.reason or ""
return d
@dataclass
class Manifest:
mode: Mode
base_ref: str
head_ref: str
default_branch: str
language_breakdown: LanguageBreakdown
changed_files: list[ChangedFile]
findings: list[Finding]
package_diffs: dict[str, PackageDiff]
tool_stats: dict[str, ToolStat]
tools_unavailable: list[str]
errors: list[str]
def to_dict(self) -> dict:
return {
"mode": self.mode,
"base_ref": self.base_ref,
"head_ref": self.head_ref,
"default_branch": self.default_branch,
"language_breakdown": self.language_breakdown.to_dict(),
"changed_files": [c.to_dict() for c in self.changed_files],
"findings": [f.to_dict() for f in self.findings],
"package_diffs": {k: v.to_dict() for k, v in self.package_diffs.items()},
"tool_stats": {k: v.to_dict() for k, v in self.tool_stats.items()},
"tools_unavailable": list(self.tools_unavailable),
"errors": list(self.errors),
}
def to_json(self, indent: int = 2) -> str:
return json.dumps(self.to_dict(), indent=indent, sort_keys=False)
@@ -0,0 +1,62 @@
"""Diff dependency manifests to produce PackageDiff."""
from __future__ import annotations
import json
import re
from scripts.manifest import PackageDiff, PackageEntry, PackageUpgrade
_REQ_LINE_RE = re.compile(
r"^\s*(?P<name>[A-Za-z0-9_\-.]+)\s*==\s*(?P<version>[^\s#]+)"
)
def _parse_requirements(text: str) -> dict[str, str]:
out: dict[str, str] = {}
for line in text.splitlines():
line = line.split("#", 1)[0]
m = _REQ_LINE_RE.match(line)
if m:
out[m.group("name").lower()] = m.group("version")
return out
def _diff_maps(before: dict[str, str], after: dict[str, str]) -> PackageDiff:
added = [
PackageEntry(name=n, version=v)
for n, v in sorted(after.items()) if n not in before
]
removed = [
PackageEntry(name=n, version=v)
for n, v in sorted(before.items()) if n not in after
]
upgraded = [
PackageUpgrade(name=n, from_version=before[n], to_version=after[n])
for n in sorted(before.keys() & after.keys())
if before[n] != after[n]
]
return PackageDiff(added=added, removed=removed, upgraded=upgraded)
def diff_requirements_txt(before: str, after: str) -> PackageDiff:
return _diff_maps(_parse_requirements(before), _parse_requirements(after))
def _parse_package_json(text: str) -> dict[str, str]:
try:
doc = json.loads(text)
except json.JSONDecodeError:
return {}
out: dict[str, str] = {}
for key in ("dependencies", "devDependencies"):
section = doc.get(key) or {}
if isinstance(section, dict):
for name, version in section.items():
if isinstance(version, str):
out[name.lower()] = version
return out
def diff_package_json(before: str, after: str) -> PackageDiff:
return _diff_maps(_parse_package_json(before), _parse_package_json(after))
@@ -0,0 +1,70 @@
"""Aggregate runs.jsonl into precision + token-cost stats."""
from __future__ import annotations
import json
import sys
from pathlib import Path
def compute_stats(log_path: Path) -> dict:
if not log_path.exists():
return {"by_agent": {}, "by_rule": {}, "runs": 0}
by_agent: dict[str, dict] = {}
by_rule: dict[str, dict] = {}
runs: set[str] = set()
for line in log_path.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
rec = json.loads(line)
agent = rec.get("agent", "?")
if rec["kind"] == "subagent_run":
runs.add(rec["run_id"])
a = by_agent.setdefault(agent, _empty_agent())
a["tokens"] += rec.get("input_tokens", 0) + rec.get("output_tokens", 0)
a["duration_ms"] += rec.get("duration_ms", 0)
a["runs"] += 1
elif rec["kind"] == "verdict":
verdict = rec["verdict"]
a = by_agent.setdefault(agent, _empty_agent())
a["total"] += 1
a[verdict] = a.get(verdict, 0) + 1
rule_key = f"{agent}/{rec['rule_id']}"
r = by_rule.setdefault(rule_key, _empty_rule())
r["total"] += 1
r[verdict] = r.get(verdict, 0) + 1
for a in by_agent.values():
a["precision"] = a["kept"] / a["total"] if a["total"] else 0.0
a["tokens_per_kept"] = a["tokens"] / a["kept"] if a["kept"] else float("inf")
for r in by_rule.values():
r["precision"] = r["kept"] / r["total"] if r["total"] else 0.0
return {"by_agent": by_agent, "by_rule": by_rule, "runs": len(runs)}
def _empty_agent() -> dict:
return {
"tokens": 0, "duration_ms": 0, "runs": 0,
"total": 0, "kept": 0, "dismissed": 0, "false_positive": 0,
}
def _empty_rule() -> dict:
return {"total": 0, "kept": 0, "dismissed": 0, "false_positive": 0}
def main(argv: list[str]) -> int:
log = Path.home() / ".claude/cache/audit-code/runs.jsonl"
stats = compute_stats(log)
print(json.dumps(stats, indent=2))
return 0
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
@@ -0,0 +1,29 @@
rules:
- id: lua-os-execute
pattern-either:
- pattern: os.execute($CMD)
- pattern: io.popen($CMD)
message: >-
Shell command executed via os.execute/io.popen. If $CMD includes any
externally-influenced data (arguments, env vars, network/file input),
this is command injection. Verify the command is a fixed literal or
properly escaped/allowlisted.
languages: [lua]
severity: WARNING
metadata:
cwe: ["CWE-78: Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection')"]
category: security
- id: lua-dynamic-code-load
pattern-either:
- pattern: load($CODE)
- pattern: loadstring($CODE)
message: >-
Dynamic code loaded via load/loadstring. If $CODE is derived from
external input, this allows arbitrary code execution. Verify the
source is trusted and not attacker-influenced.
languages: [lua]
severity: WARNING
metadata:
cwe: ["CWE-94: Improper Control of Generation of Code ('Code Injection')"]
category: security
@@ -0,0 +1,70 @@
"""Subprocess wrapper for external linters + parallel dispatch."""
from __future__ import annotations
import concurrent.futures as cf
import os
import shutil
import subprocess
from dataclasses import dataclass
from pathlib import Path
_DEFAULT_TIMEOUT = 180
_SKILL_VENV_BIN = Path(__file__).resolve().parent.parent / ".venv" / "bin"
@dataclass
class ToolResult:
binary: str
ran: bool
exit_code: int = -1
stdout: str = ""
stderr: str = ""
reason: str = ""
def resolve_tool(binary: str) -> str | None:
"""Resolve a tool name to an absolute path, preferring the skill's
`.venv/bin/` over the user's PATH. Returns None if not found anywhere.
"""
local = _SKILL_VENV_BIN / binary
if local.is_file() and os.access(local, os.X_OK):
return str(local)
return shutil.which(binary)
def tool_available(binary: str) -> bool:
return resolve_tool(binary) is not None
def run_tool(args: list[str], cwd: str, timeout: int = _DEFAULT_TIMEOUT) -> ToolResult:
binary = args[0]
resolved = resolve_tool(binary)
if resolved is None:
return ToolResult(binary=binary, ran=False, reason="not on PATH")
try:
result = subprocess.run(
[resolved, *args[1:]], cwd=cwd, capture_output=True, text=True,
check=False, timeout=timeout,
)
except subprocess.TimeoutExpired:
return ToolResult(binary=binary, ran=False, reason=f"timeout after {timeout}s")
return ToolResult(
binary=binary, ran=True,
exit_code=result.returncode,
stdout=result.stdout, stderr=result.stderr,
)
def run_tools_parallel(
jobs: list[tuple[list[str], str]], max_workers: int = 8,
) -> list[ToolResult]:
"""Run a list of (args, cwd) jobs in parallel. Order preserved."""
results: list[ToolResult] = [None] * len(jobs) # type: ignore[list-item]
with cf.ThreadPoolExecutor(max_workers=max_workers) as ex:
futures = {ex.submit(run_tool, a, c): i for i, (a, c) in enumerate(jobs)}
for fut in cf.as_completed(futures):
i = futures[fut]
results[i] = fut.result()
return results
@@ -0,0 +1,111 @@
"""Per-agent manifest slicing."""
from __future__ import annotations
_TYPE_TOOLS = {"mypy", "tsc"}
_DEP_TOOLS = {"pip-audit", "osv-scanner"}
_SECRET_TOOLS = {"gitleaks"}
_MAINTAINABILITY_TOOLS = {"vulture", "radon", "interrogate", "lizard", "knip", "jscpd", "selene"}
_GHA_TOOLS = {"actionlint", "zizmor"}
_PSSA_SECURITY_RULES = {
"PSAvoidUsingPlainTextForPassword",
"PSAvoidUsingConvertToSecureStringWithPlainText",
"PSAvoidUsingUsernameAndPasswordParams",
"PSUsePSCredentialType",
"PSAvoidUsingInvokeExpression",
"PSAvoidUsingComputerNameHardcoded",
"PSAvoidUsingBrokenHashAlgorithms",
}
def slice_for_agent(manifest: dict, agent: str) -> dict:
"""Return a subset of the manifest scoped to a specific reviewer."""
base = {
"mode": manifest["mode"],
"base_ref": manifest["base_ref"],
"head_ref": manifest["head_ref"],
"default_branch": manifest["default_branch"],
"language_breakdown": manifest["language_breakdown"],
"changed_files": list(manifest["changed_files"]),
"tool_stats": dict(manifest["tool_stats"]),
"tools_unavailable": list(manifest["tools_unavailable"]),
"errors": list(manifest["errors"]),
}
findings = manifest["findings"]
if agent == "security-triage":
kept: list[dict] = []
for f in findings:
tool = f["tool"]
if tool in {"bandit", "ruff", "opengrep", "luac", "injectionhunter"}:
kept.append(f)
elif tool == "psscriptanalyzer" and f.get("rule_id", "") in _PSSA_SECURITY_RULES:
kept.append(f)
elif tool == "eslint" and "security" in f.get("rule_id", ""):
kept.append(f)
elif tool == "dotnet" and f.get("rule_id", "").startswith("SCS"):
kept.append(f)
return {**base, "findings": kept}
if agent == "type-safety":
kept = []
for f in findings:
tool = f["tool"]
if tool in _TYPE_TOOLS:
kept.append(f)
elif tool == "eslint" and "security" not in f.get("rule_id", ""):
kept.append(f)
elif tool == "dotnet" and not f.get("rule_id", "").startswith("SCS"):
kept.append(f)
return {**base, "findings": kept}
if agent == "dependency":
return {
**base,
"findings": [f for f in findings if f["tool"] in _DEP_TOOLS],
"package_diffs": dict(manifest["package_diffs"]),
}
if agent == "secrets":
return {**base, "findings": [f for f in findings if f["tool"] in _SECRET_TOOLS]}
if agent == "consistency":
kept: list[dict] = []
for f in findings:
if f["tool"] != "ruff-idiom":
continue
rid = f.get("rule_id", "")
if rid.startswith(("C901", "PLR0915")):
continue
kept.append(f)
return {**base, "findings": kept}
if agent == "walkthrough":
return {
"mode": manifest["mode"],
"base_ref": manifest["base_ref"],
"head_ref": manifest["head_ref"],
"default_branch": manifest["default_branch"],
"language_breakdown": manifest["language_breakdown"],
"changed_files": list(manifest["changed_files"]),
"errors": list(manifest["errors"]),
}
if agent == "maintainability":
kept = []
for f in findings:
if f["tool"] in _MAINTAINABILITY_TOOLS:
kept.append(f)
elif f["tool"] == "luac":
kept.append(f)
elif f["tool"] == "psscriptanalyzer" and f.get("rule_id", "") not in _PSSA_SECURITY_RULES:
kept.append(f)
elif f["tool"] == "ruff-idiom" and f.get("rule_id", "").startswith(("C901", "PLR0915")):
kept.append(f)
return {**base, "findings": kept}
if agent == "gha-reviewer":
return {**base, "findings": [f for f in findings if f["tool"] in _GHA_TOOLS]}
return manifest
@@ -0,0 +1,53 @@
"""Append-only JSONL telemetry for audit-code runs."""
from __future__ import annotations
import json
from datetime import datetime, timezone
from pathlib import Path
def _now() -> str:
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def _append(log_path: Path, record: dict) -> None:
log_path.parent.mkdir(parents=True, exist_ok=True)
with log_path.open("a", encoding="utf-8") as f:
f.write(json.dumps(record) + "\n")
def append_subagent_run(
log_path: Path, *, run_id: str, repo: str, mode: str, agent: str,
model: str, input_tokens: int, output_tokens: int,
duration_ms: int, finding_count: int,
) -> None:
_append(log_path, {
"kind": "subagent_run", "ts": _now(),
"run_id": run_id, "repo": repo, "mode": mode, "agent": agent,
"model": model,
"input_tokens": input_tokens, "output_tokens": output_tokens,
"duration_ms": duration_ms, "finding_count": finding_count,
})
def append_verdict(
log_path: Path, *, run_id: str, agent: str, rule_id: str,
file: str, line: int, verdict: str, notes: str = "",
) -> None:
if verdict not in {"kept", "dismissed", "false_positive"}:
raise ValueError(f"invalid verdict: {verdict!r}")
_append(log_path, {
"kind": "verdict", "ts": _now(),
"run_id": run_id, "agent": agent, "rule_id": rule_id,
"file": file, "line": line, "verdict": verdict, "notes": notes,
})
def read_runs(log_path: Path) -> list[dict]:
if not log_path.exists():
return []
return [
json.loads(line)
for line in log_path.read_text(encoding="utf-8").splitlines()
if line.strip()
]
@@ -0,0 +1,5 @@
import sys
from pathlib import Path
_root = Path(__file__).resolve().parent.parent
sys.path.insert(0, str(_root))
@@ -0,0 +1,18 @@
[
{
"message": "shellcheck reported issue in this script: SC2086:info:1:6: Double quote to prevent globbing and word splitting",
"filepath": ".github/workflows/ci.yml",
"line": 22,
"column": 9,
"kind": "shellcheck",
"snippet": " run: echo $SOME_VAR"
},
{
"message": "property \"foo\" is not defined in object type {}",
"filepath": ".github/workflows/ci.yml",
"line": 15,
"column": 20,
"kind": "expression",
"snippet": " - run: echo ${{ foo }}"
}
]
@@ -0,0 +1,28 @@
{
"results": [
{
"code": "...",
"filename": "src/api.py",
"issue_confidence": "HIGH",
"issue_severity": "HIGH",
"issue_text": "Possible SQL injection vector through string-based query construction.",
"line_number": 45,
"line_range": [45, 45],
"more_info": "https://bandit.readthedocs.io/en/latest/plugins/b608_hardcoded_sql_expressions.html",
"test_id": "B608",
"test_name": "hardcoded_sql_expressions"
},
{
"code": "...",
"filename": "src/api.py",
"issue_confidence": "MEDIUM",
"issue_severity": "LOW",
"issue_text": "Use of assert detected.",
"line_number": 12,
"line_range": [12, 12],
"test_id": "B101",
"test_name": "assert_used"
}
],
"metrics": {"_totals": {"SEVERITY.HIGH": 1, "SEVERITY.LOW": 1}}
}
@@ -0,0 +1,10 @@
Microsoft (R) Build Engine version 17.0
Copyright (C) Microsoft Corporation. All rights reserved.
/repo/src/Program.cs(22,5): warning SCS0018: Path traversal: injection possible. [/repo/src/MyApp.csproj]
/repo/src/Program.cs(45,12): error CS8602: Dereference of a possibly null reference. [/repo/src/MyApp.csproj]
/repo/src/Program.cs(60,8): warning CA1834: Use 'StringBuilder.Append(char)' for single-character strings. [/repo/src/MyApp.csproj]
Build succeeded.
1 Warning(s)
0 Error(s)
@@ -0,0 +1,33 @@
[
{
"filePath": "/repo/src/auth.ts",
"messages": [
{
"ruleId": "security/detect-eval-with-expression",
"severity": 2,
"message": "eval with non-literal expression",
"line": 22,
"column": 5,
"endLine": 22,
"endColumn": 35
},
{
"ruleId": "@typescript-eslint/no-explicit-any",
"severity": 1,
"message": "Unexpected any. Specify a different type.",
"line": 8,
"column": 12,
"endLine": 8,
"endColumn": 15
}
],
"errorCount": 1,
"warningCount": 1
},
{
"filePath": "/repo/src/utils.ts",
"messages": [],
"errorCount": 0,
"warningCount": 0
}
]
@@ -0,0 +1,20 @@
[
{
"RuleID": "aws-access-token",
"Description": "AWS Access Token",
"StartLine": 12,
"EndLine": 12,
"File": "src/config.py",
"Secret": "redacted",
"Match": "aws_access_key_id assignment"
},
{
"RuleID": "generic-api-key",
"Description": "Generic API Key",
"StartLine": 30,
"EndLine": 30,
"File": "src/api.py",
"Secret": "redacted",
"Match": "API key assignment"
}
]
@@ -0,0 +1,18 @@
[
{
"RuleName": "InjectionRisk.InvokeExpression",
"Severity": "Warning",
"ScriptPath": "/repo/scripts/Deploy.ps1",
"Line": 12,
"EndLine": 12,
"Message": "Possible script injection risk via the Invoke-Expression cmdlet."
},
{
"RuleName": "InjectionRisk.AddScriptBlock",
"Severity": "Warning",
"ScriptPath": "/repo/scripts/Deploy.ps1",
"Line": 33,
"EndLine": 35,
"Message": "Possible script injection risk via the AddScript method."
}
]
@@ -0,0 +1,5 @@
src/api.py:12: error: Incompatible types in assignment (expression has type "str", variable has type "int") [assignment]
src/api.py:45: error: Argument 1 has incompatible type "Any"; expected "str" [arg-type]
src/utils.py:8: note: Possible overload variants:
src/api.py:60: warning: Unused "type: ignore" comment [unused-ignore]
Found 3 errors in 2 files (checked 5 source files)
@@ -0,0 +1,27 @@
{
"version": "1.25.0",
"results": [
{
"check_id": "python.lang.security.audit.dangerous-code.dangerous-code",
"path": "src/api.py",
"start": {"line": 12, "col": 5},
"end": {"line": 12, "col": 30},
"extra": {
"severity": "ERROR",
"message": "Detected use of dynamic code construction.",
"metadata": {"cwe": ["CWE-94: Improper Control of Generation of Code ('Code Injection')"]}
}
},
{
"check_id": "javascript.lang.audit.dynamic-expression",
"path": "src/auth.ts",
"start": {"line": 22, "col": 5},
"end": {"line": 22, "col": 35},
"extra": {
"severity": "WARNING",
"message": "Dynamic expression with non-literal input",
"metadata": {"cwe": ["CWE-95"]}
}
}
]
}
@@ -0,0 +1,18 @@
{
"results": [
{
"source": {"path": "package-lock.json", "type": "lockfile"},
"packages": [
{
"package": {"name": "axios", "version": "0.21.0", "ecosystem": "npm"},
"vulnerabilities": [
{
"id": "GHSA-cph5-m8f7-6c5x",
"summary": "Axios is vulnerable to Server-Side Request Forgery"
}
]
}
]
}
]
}
@@ -0,0 +1,20 @@
{
"dependencies": [
{
"name": "pyyaml",
"version": "5.4.1",
"vulns": [
{
"id": "GHSA-8q59-q68h-6hv4",
"fix_versions": ["6.0"],
"description": "PyYAML vulnerable to arbitrary code generation"
}
]
},
{
"name": "requests",
"version": "2.31.0",
"vulns": []
}
]
}
@@ -0,0 +1,42 @@
[
{
"RuleName": "PSAvoidUsingInvokeExpression",
"Severity": "Warning",
"ScriptPath": "/repo/scripts/Deploy.ps1",
"Line": 12,
"EndLine": 12,
"Message": "Invoke-Expression is used. Please remove Invoke-Expression from script and find other options instead."
},
{
"RuleName": "PSAvoidUsingConvertToSecureStringWithPlainText",
"Severity": "Error",
"ScriptPath": "/repo/scripts/Deploy.ps1",
"Line": 20,
"EndLine": 22,
"Message": "File 'Deploy.ps1' uses ConvertTo-SecureString with plaintext."
},
{
"RuleName": "PSAvoidUsingWriteHost",
"Severity": "Warning",
"ScriptPath": "/repo/module/Helpers.psm1",
"Line": 5,
"EndLine": 5,
"Message": "File 'Helpers.psm1' uses Write-Host."
},
{
"RuleName": "PSUseDeclaredVarsMoreThanAssignments",
"Severity": "Information",
"ScriptPath": "/repo/module/Helpers.psm1",
"Line": 30,
"EndLine": 30,
"Message": "The variable 'unused' is assigned but never used."
},
{
"RuleName": "TypeNotFound",
"Severity": "ParseError",
"ScriptPath": "/repo/module/Helpers.psm1",
"Line": 41,
"EndLine": 41,
"Message": "Unable to find type [Foo]."
}
]
@@ -0,0 +1,23 @@
[
{
"code": "S608",
"location": {"row": 45, "column": 12},
"end_location": {"row": 45, "column": 80},
"filename": "src/api.py",
"message": "Possible SQL injection vector through string-based query construction"
},
{
"code": "E501",
"location": {"row": 30, "column": 1},
"end_location": {"row": 30, "column": 105},
"filename": "src/api.py",
"message": "Line too long (104 > 88)"
},
{
"code": "F401",
"location": {"row": 3, "column": 1},
"end_location": {"row": 3, "column": 20},
"filename": "src/api.py",
"message": "imported but unused"
}
]
@@ -0,0 +1,44 @@
[
{
"filename": "src/game.lua",
"primary_label": "undefined variable `undefined_var`",
"secondary_labels": [],
"notes": [],
"code": {
"name": "undefined_variable",
"severity": "Error"
},
"span": {
"start": {"line": 5, "character": 1},
"end": {"line": 5, "character": 14}
}
},
{
"filename": "src/game.lua",
"primary_label": "accessing undefined field `nonexistent` on type `table`",
"secondary_labels": [],
"notes": [],
"code": {
"name": "undefined_field",
"severity": "Warning"
},
"span": {
"start": {"line": 12, "character": 3},
"end": {"line": 12, "character": 15}
}
},
{
"filename": "src/game.lua",
"primary_label": "variable `_unused` is set but never used",
"secondary_labels": [],
"notes": [],
"code": {
"name": "unused_variable",
"severity": "Note"
},
"span": {
"start": {"line": 20, "character": 5},
"end": {"line": 20, "character": 12}
}
}
]
@@ -0,0 +1,3 @@
src/auth.ts(22,5): error TS2322: Type 'string' is not assignable to type 'number'.
src/utils.ts(8,12): error TS2345: Argument of type 'unknown' is not assignable to parameter of type 'string'.
src/index.ts(3,1): error TS6133: 'fs' is declared but its value is never read.
@@ -0,0 +1,65 @@
{
"version": "2.1.0",
"runs": [
{
"tool": {
"driver": {
"name": "zizmor",
"version": "1.0.0",
"rules": [
{
"id": "unpinned-uses",
"name": "Unpinned uses",
"shortDescription": {"text": "Action ref is not pinned to a SHA digest"},
"properties": {"tags": ["CWE-829"]}
},
{
"id": "excessive-permissions",
"name": "Excessive permissions",
"shortDescription": {"text": "Workflow grants write-all permissions"},
"properties": {"tags": []}
}
]
}
},
"results": [
{
"ruleId": "unpinned-uses",
"level": "warning",
"message": {"text": "uses: actions/checkout@v4 is not pinned by digest"},
"locations": [
{
"physicalLocation": {
"artifactLocation": {"uri": ".github/workflows/ci.yml"},
"region": {
"startLine": 12,
"startColumn": 9,
"endLine": 12,
"endColumn": 38
}
}
}
]
},
{
"ruleId": "excessive-permissions",
"level": "error",
"message": {"text": "Workflow uses write-all permissions; scope to minimum required"},
"locations": [
{
"physicalLocation": {
"artifactLocation": {"uri": ".github/workflows/ci.yml"},
"region": {
"startLine": 5,
"startColumn": 1,
"endLine": 5,
"endColumn": 30
}
}
}
]
}
]
}
]
}
@@ -0,0 +1,41 @@
from pathlib import Path
from scripts.adapters.actionlint import parse_actionlint_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_findings():
payload = (_FIXTURES / "actionlint_sample.json").read_text()
findings = parse_actionlint_output(payload)
assert len(findings) == 2
def test_shellcheck_finding_fields():
payload = (_FIXTURES / "actionlint_sample.json").read_text()
findings = parse_actionlint_output(payload)
sc = next(f for f in findings if f.rule_id == "shellcheck")
assert sc.tool == "actionlint"
assert sc.severity == "high"
assert sc.file == ".github/workflows/ci.yml"
assert sc.line == 22
assert sc.end_line == 22
assert "SC2086" in sc.message
def test_expression_finding_fields():
payload = (_FIXTURES / "actionlint_sample.json").read_text()
findings = parse_actionlint_output(payload)
expr = next(f for f in findings if f.rule_id == "expression")
assert expr.line == 15
def test_handles_empty_array():
findings = parse_actionlint_output("[]")
assert findings == []
def test_handles_invalid_json():
findings = parse_actionlint_output("not json")
assert findings == []
@@ -0,0 +1,28 @@
from pathlib import Path
from scripts.adapters.bandit import parse_bandit_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_findings_with_severity_lowercase():
payload = (_FIXTURES / "bandit_sample.json").read_text()
findings = parse_bandit_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert "B608" in by_rule
assert by_rule["B608"].severity == "high"
assert by_rule["B608"].file == "src/api.py"
assert by_rule["B608"].line == 45
assert by_rule["B608"].end_line == 45
assert "SQL injection" in by_rule["B608"].message
def test_handles_empty_results():
findings = parse_bandit_output('{"results": [], "metrics": {}}')
assert findings == []
def test_handles_invalid_json():
findings = parse_bandit_output('not json')
assert findings == []
@@ -0,0 +1,34 @@
from pathlib import Path
from scripts.adapters.dotnet import parse_dotnet_build_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_scs_as_high():
txt = (_FIXTURES / "dotnet_build_sample.txt").read_text()
findings = parse_dotnet_build_output(txt, repo_root="/repo")
by_rule = {f.rule_id: f for f in findings}
assert "SCS0018" in by_rule
assert by_rule["SCS0018"].severity == "high"
assert by_rule["SCS0018"].file == "src/Program.cs"
assert by_rule["SCS0018"].line == 22
def test_parses_cs_error_as_medium():
txt = (_FIXTURES / "dotnet_build_sample.txt").read_text()
findings = parse_dotnet_build_output(txt, repo_root="/repo")
by_rule = {f.rule_id: f for f in findings}
assert by_rule["CS8602"].severity == "medium"
def test_parses_ca_warning_as_low():
txt = (_FIXTURES / "dotnet_build_sample.txt").read_text()
findings = parse_dotnet_build_output(txt, repo_root="/repo")
by_rule = {f.rule_id: f for f in findings}
assert by_rule["CA1834"].severity == "low"
def test_handles_empty():
assert parse_dotnet_build_output("", repo_root="/repo") == []
@@ -0,0 +1,41 @@
from pathlib import Path
from scripts.adapters.eslint import parse_eslint_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_security_rule_as_high():
payload = (_FIXTURES / "eslint_sample.json").read_text()
findings = parse_eslint_output(payload, repo_root="/repo")
rules = {f.rule_id: f for f in findings}
assert "security/detect-eval-with-expression" in rules
sec = rules["security/detect-eval-with-expression"]
assert sec.severity == "high"
assert sec.file == "src/auth.ts"
assert sec.line == 22
def test_warning_severity_low():
payload = (_FIXTURES / "eslint_sample.json").read_text()
findings = parse_eslint_output(payload, repo_root="/repo")
rules = {f.rule_id: f for f in findings}
assert rules["@typescript-eslint/no-explicit-any"].severity == "low"
def test_non_security_error_severity_medium():
payload = '[{"filePath": "/repo/a.js", "messages": [{"ruleId": "no-undef", "severity": 2, "message": "x", "line": 1, "column": 1, "endLine": 1, "endColumn": 5}]}]'
findings = parse_eslint_output(payload, repo_root="/repo")
assert findings[0].severity == "medium"
def test_strips_repo_prefix():
payload = (_FIXTURES / "eslint_sample.json").read_text()
findings = parse_eslint_output(payload, repo_root="/repo")
for f in findings:
assert not f.file.startswith("/")
def test_handles_invalid_json():
assert parse_eslint_output("nope", repo_root="/repo") == []
@@ -0,0 +1,25 @@
from pathlib import Path
from scripts.adapters.gitleaks import parse_gitleaks_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_findings_as_critical():
payload = (_FIXTURES / "gitleaks_sample.json").read_text()
findings = parse_gitleaks_output(payload)
rules = {f.rule_id for f in findings}
assert "aws-access-token" in rules
assert "generic-api-key" in rules
for f in findings:
assert f.severity == "critical"
assert f.tool == "gitleaks"
def test_handles_empty():
assert parse_gitleaks_output("[]") == []
def test_handles_invalid():
assert parse_gitleaks_output("nope") == []
@@ -0,0 +1,69 @@
import json
from scripts.adapters.interrogate import parse_interrogate_output
def test_parses_missing_public_docstring():
payload = {
"files": {
"src/foo.py": {
"missing": [
{"name": "src/foo.py:my_public_fn", "type": "function", "lineno": 12, "private": False}
]
}
}
}
findings = parse_interrogate_output(json.dumps(payload))
assert len(findings) == 1
f = findings[0]
assert f.tool == "interrogate"
assert f.file == "src/foo.py"
assert f.line == 12
assert "my_public_fn" in f.message
assert f.rule_id == "interrogate:missing-docstring"
def test_skips_private_symbol():
payload = {
"files": {
"src/foo.py": {
"missing": [
{"name": "src/foo.py:_helper", "type": "function", "lineno": 30, "private": True}
]
}
}
}
assert parse_interrogate_output(json.dumps(payload)) == []
def test_skips_underscore_named_symbol_even_if_private_false():
payload = {
"files": {
"src/foo.py": {
"missing": [
{"name": "src/foo.py:_internal", "type": "function", "lineno": 5, "private": False}
]
}
}
}
assert parse_interrogate_output(json.dumps(payload)) == []
def test_low_severity_always():
payload = {
"files": {
"src/foo.py": {
"missing": [
{"name": "src/foo.py:pub_fn", "type": "function", "lineno": 1, "private": False}
]
}
}
}
findings = parse_interrogate_output(json.dumps(payload))
assert findings[0].severity == "low"
def test_empty_or_malformed_input():
assert parse_interrogate_output("") == []
assert parse_interrogate_output("not json") == []
assert parse_interrogate_output("{}") == []
@@ -0,0 +1,47 @@
import json
from scripts.adapters.jscpd import parse_jscpd_output
def test_parses_small_clone_as_low():
payload = {
"duplicates": [
{
"firstFile": {"name": "src/a.ts", "start": 1, "end": 10},
"secondFile": {"name": "src/b.ts", "start": 5, "end": 14},
"lines": 10,
}
]
}
findings = parse_jscpd_output(json.dumps(payload))
assert len(findings) == 1
f = findings[0]
assert f.tool == "jscpd"
assert f.severity == "low"
assert f.rule_id == "jscpd:clone-10lines"
assert f.file == "src/a.ts"
assert f.line == 1
assert f.end_line == 10
assert "src/b.ts" in f.message
assert "10 lines" in f.message
def test_parses_large_clone_as_medium():
payload = {
"duplicates": [
{
"firstFile": {"name": "src/a.ts", "start": 1, "end": 40},
"secondFile": {"name": "src/b.ts", "start": 1, "end": 40},
"lines": 40,
}
]
}
findings = parse_jscpd_output(json.dumps(payload))
assert findings[0].severity == "medium"
def test_empty_or_no_duplicates():
assert parse_jscpd_output("") == []
assert parse_jscpd_output("{}") == []
assert parse_jscpd_output(json.dumps({"duplicates": []})) == []
assert parse_jscpd_output("not json") == []
@@ -0,0 +1,46 @@
import json
from scripts.adapters.knip import parse_knip_output
def test_parses_dead_export():
payload = {
"issues": [
{
"file": "src/foo.ts",
"exports": [{"name": "unusedExport", "line": 12, "col": 1}],
}
]
}
findings = parse_knip_output(json.dumps(payload))
assert len(findings) == 1
f = findings[0]
assert f.tool == "knip"
assert f.rule_id == "knip:dead-export"
assert f.severity == "medium"
assert f.file == "src/foo.ts"
assert f.line == 12
assert "unusedExport" in f.message
def test_multiple_exports():
payload = {
"issues": [
{
"file": "src/foo.ts",
"exports": [
{"name": "exportA", "line": 5, "col": 1},
{"name": "exportB", "line": 10, "col": 1},
],
}
]
}
findings = parse_knip_output(json.dumps(payload))
assert len(findings) == 2
def test_empty_or_no_issues():
assert parse_knip_output("") == []
assert parse_knip_output("{}") == []
assert parse_knip_output(json.dumps({"issues": []})) == []
assert parse_knip_output("not json") == []
@@ -0,0 +1,35 @@
from scripts.adapters.lizard import parse_lizard_output
def test_parses_high_ccn():
row = "12,15,80,3,15,my_fn@10@src/foo.py\n"
findings = parse_lizard_output(row)
assert len(findings) == 1
f = findings[0]
assert f.tool == "lizard"
assert f.severity == "medium"
assert f.rule_id == "lizard:ccn=15"
assert f.file == "src/foo.py"
assert f.line == 10
assert f.end_line == 10
assert "my_fn" in f.message
assert "15" in f.message
def test_very_high_ccn_is_high():
row = "30,25,200,5,40,complex_fn@5@src/bar.py\n"
findings = parse_lizard_output(row)
assert findings[0].severity == "high"
def test_low_ccn_skipped():
row = "5,3,30,1,6,simple@1@src/foo.py\n"
assert parse_lizard_output(row) == []
def test_malformed_row_skipped():
assert parse_lizard_output("1,2,3\n") == []
def test_empty_input():
assert parse_lizard_output("") == []
@@ -0,0 +1,43 @@
from scripts.adapters.luac import parse_luac_output
def test_parses_single_syntax_error():
stderr = "luac: ./src/game.lua:8: 'end' expected (to close 'function' at line 1) near '<eof>'"
findings = parse_luac_output(stderr, repo_root="")
assert len(findings) == 1
f = findings[0]
assert f.tool == "luac"
assert f.rule_id == "syntax-error"
assert f.severity == "critical"
assert f.line == 8
assert f.end_line == 8
assert "end" in f.message
def test_parses_multiple_errors():
stderr = (
"luac: ./a.lua:3: unexpected symbol near '='\n"
"luac: ./b.lua:17: 'end' expected near '<eof>'"
)
findings = parse_luac_output(stderr, repo_root="")
assert len(findings) == 2
files = {f.file for f in findings}
assert "./a.lua" in files
assert "./b.lua" in files
def test_strips_repo_root_prefix():
stderr = "luac: /home/user/myrepo/src/game.lua:5: unexpected symbol near '+'"
findings = parse_luac_output(stderr, repo_root="/home/user/myrepo")
assert findings[0].file == "src/game.lua"
def test_empty_stderr_returns_empty():
findings = parse_luac_output("", repo_root="")
assert findings == []
def test_non_matching_lines_ignored():
stderr = "warning: something else\nnot a luac error"
findings = parse_luac_output(stderr, repo_root="")
assert findings == []
@@ -0,0 +1,32 @@
from pathlib import Path
from scripts.adapters.mypy import parse_mypy_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_error_lines():
txt = (_FIXTURES / "mypy_sample.txt").read_text()
findings = parse_mypy_output(txt)
by_line = {f.line: f for f in findings if f.file == "src/api.py"}
assert 12 in by_line
assert by_line[12].severity == "medium"
assert by_line[12].rule_id == "assignment"
def test_parses_warning_lines():
txt = (_FIXTURES / "mypy_sample.txt").read_text()
findings = parse_mypy_output(txt)
warns = [f for f in findings if f.severity == "low"]
assert any(f.line == 60 and f.rule_id == "unused-ignore" for f in warns)
def test_skips_note_lines():
txt = (_FIXTURES / "mypy_sample.txt").read_text()
findings = parse_mypy_output(txt)
assert not any(f.file == "src/utils.py" and f.line == 8 for f in findings)
def test_handles_empty_input():
assert parse_mypy_output("") == []
@@ -0,0 +1,43 @@
from pathlib import Path
from scripts.adapters.opengrep import parse_opengrep_output, _clean_rule_id
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_error_as_high():
payload = (_FIXTURES / "opengrep_sample.json").read_text()
findings = parse_opengrep_output(payload)
by_rule = {f.rule_id: f for f in findings}
key = "python.lang.security.audit.dangerous-code.dangerous-code"
assert key in by_rule
assert by_rule[key].severity == "high"
assert by_rule[key].cwe == "CWE-94"
assert by_rule[key].line == 12
def test_parses_warning_as_medium():
payload = (_FIXTURES / "opengrep_sample.json").read_text()
findings = parse_opengrep_output(payload)
by_rule = {f.rule_id: f for f in findings}
key = "javascript.lang.audit.dynamic-expression"
assert by_rule[key].severity == "medium"
def test_handles_empty():
assert parse_opengrep_output('{"results": []}') == []
def test_handles_invalid():
assert parse_opengrep_output("nope") == []
def test_clean_rule_id_strips_local_rules_path_prefix():
mangled = "home.mroberts..claude.plugins.reviews.skills.audit-code.scripts.rules.lua-os-execute"
assert _clean_rule_id(mangled) == "lua-os-execute"
def test_clean_rule_id_leaves_registry_ids_untouched():
registry_id = "python.lang.security.audit.dangerous-code.dangerous-code"
assert _clean_rule_id(registry_id) == registry_id
@@ -0,0 +1,26 @@
from pathlib import Path
from scripts.adapters.osv_scanner import parse_osv_scanner_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_emits_finding_per_vuln():
payload = (_FIXTURES / "osv_scanner_sample.json").read_text()
findings = parse_osv_scanner_output(payload)
assert len(findings) == 1
f = findings[0]
assert f.rule_id == "GHSA-cph5-m8f7-6c5x"
assert f.tool == "osv-scanner"
assert f.severity == "high"
assert "axios" in f.message
assert f.file == "package-lock.json"
def test_handles_empty():
assert parse_osv_scanner_output('{"results": []}') == []
def test_handles_invalid():
assert parse_osv_scanner_output("nope") == []
@@ -0,0 +1,27 @@
from pathlib import Path
from scripts.adapters.pip_audit import parse_pip_audit_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_emits_finding_per_vuln():
payload = (_FIXTURES / "pip_audit_sample.json").read_text()
findings = parse_pip_audit_output(payload, manifest_path="requirements.txt")
assert len(findings) == 1
f = findings[0]
assert f.rule_id == "GHSA-8q59-q68h-6hv4"
assert f.severity == "high"
assert f.tool == "pip-audit"
assert "pyyaml" in f.message
assert f.file == "requirements.txt"
def test_no_finding_when_no_vulns():
payload = '{"dependencies": [{"name": "x", "version": "1", "vulns": []}]}'
assert parse_pip_audit_output(payload, manifest_path="requirements.txt") == []
def test_handles_invalid():
assert parse_pip_audit_output("nope", manifest_path="requirements.txt") == []
@@ -0,0 +1,78 @@
from pathlib import Path
from scripts.adapters.psscriptanalyzer import parse_psscriptanalyzer_output
_FIXTURES = Path(__file__).parent / "fixtures"
_REPO = "/repo"
def _pssa():
payload = (_FIXTURES / "psscriptanalyzer_sample.json").read_text()
return {f.rule_id: f for f in parse_psscriptanalyzer_output(payload, repo_root=_REPO)}
def _injection():
payload = (_FIXTURES / "injectionhunter_sample.json").read_text()
return {
f.rule_id: f
for f in parse_psscriptanalyzer_output(payload, repo_root=_REPO, tool="injectionhunter")
}
def test_paths_are_relative_to_repo_root():
f = _pssa()["PSAvoidUsingInvokeExpression"]
assert f.file == "scripts/Deploy.ps1"
assert f.tool == "psscriptanalyzer"
def test_severity_mapping():
by_rule = _pssa()
assert by_rule["PSAvoidUsingConvertToSecureStringWithPlainText"].severity == "high"
assert by_rule["PSAvoidUsingInvokeExpression"].severity == "medium"
assert by_rule["PSUseDeclaredVarsMoreThanAssignments"].severity == "low"
assert by_rule["TypeNotFound"].severity == "critical"
def test_line_range_from_extent():
f = _pssa()["PSAvoidUsingConvertToSecureStringWithPlainText"]
assert f.line == 20
assert f.end_line == 22
def test_message_preserved():
assert "Invoke-Expression" in _pssa()["PSAvoidUsingInvokeExpression"].message
def test_injectionhunter_findings_floor_at_high():
by_rule = _injection()
assert by_rule["InjectionRisk.InvokeExpression"].severity == "high"
assert by_rule["InjectionRisk.AddScriptBlock"].severity == "high"
def test_injectionhunter_tool_name_and_location():
f = _injection()["InjectionRisk.AddScriptBlock"]
assert f.tool == "injectionhunter"
assert f.file == "scripts/Deploy.ps1"
assert f.line == 33
assert f.end_line == 35
def test_single_object_not_array():
payload = '{"RuleName":"PSAvoidUsingWriteHost","Severity":"Warning","ScriptPath":"/repo/a.ps1","Line":3,"EndLine":3,"Message":"m"}'
findings = parse_psscriptanalyzer_output(payload, repo_root=_REPO)
assert len(findings) == 1
assert findings[0].file == "a.ps1"
def test_handles_empty_array():
assert parse_psscriptanalyzer_output("[]", repo_root=_REPO) == []
def test_handles_invalid_json():
assert parse_psscriptanalyzer_output("not json", repo_root=_REPO) == []
def test_unknown_severity_defaults_to_medium():
payload = '[{"RuleName":"X","Severity":"Bogus","ScriptPath":"/repo/a.ps1","Line":1,"EndLine":1,"Message":"m"}]'
assert parse_psscriptanalyzer_output(payload, repo_root=_REPO)[0].severity == "medium"
@@ -0,0 +1,58 @@
import json
from scripts.adapters.radon import parse_radon_output
def test_parses_high_complexity_as_medium():
payload = {
"src/foo.py": [
{"type": "function", "name": "bar", "lineno": 10, "endline": 25, "complexity": 12, "rank": "C"}
]
}
findings = parse_radon_output(json.dumps(payload))
assert len(findings) == 1
f = findings[0]
assert f.severity == "medium"
assert f.rule_id == "radon:cc=12"
assert f.file == "src/foo.py"
assert f.line == 10
assert f.end_line == 25
def test_very_high_complexity_is_high():
payload = {
"src/foo.py": [
{"type": "function", "name": "baz", "lineno": 30, "endline": 80, "complexity": 22, "rank": "E"}
]
}
findings = parse_radon_output(json.dumps(payload))
assert findings[0].severity == "high"
def test_low_complexity_is_skipped():
payload = {
"src/foo.py": [
{"type": "function", "name": "simple", "lineno": 1, "endline": 5, "complexity": 5, "rank": "A"}
]
}
assert parse_radon_output(json.dumps(payload)) == []
def test_multiple_files():
payload = {
"src/a.py": [
{"type": "function", "name": "fn_a", "lineno": 1, "endline": 20, "complexity": 15, "rank": "C"}
],
"src/b.py": [
{"type": "function", "name": "fn_b", "lineno": 5, "endline": 50, "complexity": 25, "rank": "E"}
],
}
findings = parse_radon_output(json.dumps(payload))
assert len(findings) == 2
files = {f.file for f in findings}
assert files == {"src/a.py", "src/b.py"}
def test_empty_input():
assert parse_radon_output("") == []
assert parse_radon_output("not json") == []
@@ -0,0 +1,33 @@
from pathlib import Path
from scripts.adapters.ruff import parse_ruff_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_findings():
payload = (_FIXTURES / "ruff_sample.json").read_text()
findings = parse_ruff_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert "S608" in by_rule
assert by_rule["S608"].severity == "high"
assert by_rule["S608"].file == "src/api.py"
assert by_rule["S608"].line == 45
def test_severity_inferred_from_prefix():
payload = (_FIXTURES / "ruff_sample.json").read_text()
findings = parse_ruff_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert by_rule["S608"].severity == "high"
assert by_rule["E501"].severity == "low"
assert by_rule["F401"].severity == "low"
def test_handles_empty():
assert parse_ruff_output("[]") == []
def test_handles_invalid_json():
assert parse_ruff_output("not json") == []
@@ -0,0 +1,35 @@
"""Parser-level tests for the idiom ruff adapter."""
import json
from scripts.adapters.ruff_idiom import parse_ruff_idiom_output
def _ruff_item(code: str, file: str = "foo.py", row: int = 10) -> dict:
return {
"code": code,
"filename": file,
"location": {"row": row, "column": 1},
"end_location": {"row": row, "column": 1},
"message": f"{code} suggestion",
}
def test_idiom_severity_mapping():
payload = json.dumps([
_ruff_item("SIM117"),
_ruff_item("UP008"),
_ruff_item("PLR0913"),
_ruff_item("C901"),
_ruff_item("B008"),
])
findings = parse_ruff_idiom_output(payload)
sev = {f.rule_id: f.severity for f in findings}
assert sev["SIM117"] == "low"
assert sev["UP008"] == "low"
assert sev["PLR0913"] == "medium"
assert sev["C901"] == "medium"
assert sev["B008"] == "medium"
def test_idiom_handles_empty_stdout():
assert parse_ruff_idiom_output("") == []
assert parse_ruff_idiom_output("not json") == []
@@ -0,0 +1,51 @@
from pathlib import Path
from scripts.adapters.selene import parse_selene_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_error_as_high():
payload = (_FIXTURES / "selene_sample.json").read_text()
findings = parse_selene_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert "undefined_variable" in by_rule
f = by_rule["undefined_variable"]
assert f.severity == "high"
assert f.file == "src/game.lua"
assert f.line == 5
assert f.end_line == 5
assert "undefined_var" in f.message
assert f.tool == "selene"
def test_parses_warning_as_medium():
payload = (_FIXTURES / "selene_sample.json").read_text()
findings = parse_selene_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert by_rule["undefined_field"].severity == "medium"
assert by_rule["undefined_field"].line == 12
assert by_rule["undefined_field"].end_line == 12
def test_parses_note_as_low():
payload = (_FIXTURES / "selene_sample.json").read_text()
findings = parse_selene_output(payload)
by_rule = {f.rule_id: f for f in findings}
assert by_rule["unused_variable"].severity == "low"
def test_handles_empty_array():
findings = parse_selene_output("[]")
assert findings == []
def test_handles_invalid_json():
findings = parse_selene_output("not json")
assert findings == []
def test_handles_non_array_json():
findings = parse_selene_output('{"error": "unexpected"}')
assert findings == []
@@ -0,0 +1,24 @@
from pathlib import Path
from scripts.adapters.tsc import parse_tsc_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_error_lines():
txt = (_FIXTURES / "tsc_sample.txt").read_text()
findings = parse_tsc_output(txt)
by_file = {f.file: f for f in findings}
assert "src/auth.ts" in by_file
assert by_file["src/auth.ts"].line == 22
assert by_file["src/auth.ts"].rule_id == "TS2322"
assert by_file["src/auth.ts"].severity == "medium"
def test_handles_empty():
assert parse_tsc_output("") == []
def test_handles_malformed_lines():
assert parse_tsc_output("not a tsc line\n") == []
@@ -0,0 +1,33 @@
from scripts.adapters.vulture import parse_vulture_output
def test_parses_unused_function():
out = "src/foo.py:42: unused function 'bar' (60% confidence)\n"
findings = parse_vulture_output(out)
assert len(findings) == 1
f = findings[0]
assert f.tool == "vulture"
assert f.file == "src/foo.py"
assert f.line == 42
assert "unused function 'bar'" in f.message
assert f.severity == "low"
def test_high_confidence_unused_is_medium():
out = "src/foo.py:7: unused import 'os' (90% confidence)\n"
findings = parse_vulture_output(out)
assert findings[0].severity == "medium"
def test_empty_stdout():
assert parse_vulture_output("") == []
def test_rule_id_contains_confidence():
out = "src/foo.py:7: unused import 'os' (90% confidence)\n"
findings = parse_vulture_output(out)
assert findings[0].rule_id == "vulture:90pct"
def test_malformed_line_skipped():
assert parse_vulture_output("not a valid line\n") == []
@@ -0,0 +1,56 @@
from pathlib import Path
from scripts.adapters.zizmor import parse_zizmor_output
_FIXTURES = Path(__file__).parent / "fixtures"
def test_parses_two_results():
payload = (_FIXTURES / "zizmor_sample.json").read_text()
findings = parse_zizmor_output(payload)
assert len(findings) == 2
def test_warning_maps_to_medium():
payload = (_FIXTURES / "zizmor_sample.json").read_text()
findings = parse_zizmor_output(payload)
unpinned = next(f for f in findings if f.rule_id == "unpinned-uses")
assert unpinned.tool == "zizmor"
assert unpinned.severity == "medium"
assert unpinned.file == ".github/workflows/ci.yml"
assert unpinned.line == 12
assert unpinned.end_line == 12
assert "actions/checkout" in unpinned.message
def test_error_maps_to_high():
payload = (_FIXTURES / "zizmor_sample.json").read_text()
findings = parse_zizmor_output(payload)
perms = next(f for f in findings if f.rule_id == "excessive-permissions")
assert perms.severity == "high"
assert perms.line == 5
def test_cwe_extracted_from_rule_tags():
payload = (_FIXTURES / "zizmor_sample.json").read_text()
findings = parse_zizmor_output(payload)
unpinned = next(f for f in findings if f.rule_id == "unpinned-uses")
assert unpinned.cwe == "CWE-829"
def test_no_cwe_when_tags_empty():
payload = (_FIXTURES / "zizmor_sample.json").read_text()
findings = parse_zizmor_output(payload)
perms = next(f for f in findings if f.rule_id == "excessive-permissions")
assert perms.cwe is None
def test_handles_empty_runs():
findings = parse_zizmor_output('{"version": "2.1.0", "runs": []}')
assert findings == []
def test_handles_invalid_json():
findings = parse_zizmor_output("not json")
assert findings == []
@@ -0,0 +1,195 @@
import importlib.util
import json
from pathlib import Path
from unittest.mock import patch, MagicMock
_SKILL_ROOT = Path(__file__).resolve().parent.parent
def _load_cli():
spec = importlib.util.spec_from_file_location(
"collect_findings", _SKILL_ROOT / "scripts" / "collect-findings.py"
)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def _fake_subprocess(cmd, **kwargs):
if cmd[:2] == ["git", "-C"]:
sub = cmd[2:]
else:
sub = cmd
if "symbolic-ref" in sub:
return MagicMock(returncode=0, stdout="refs/remotes/origin/main\n", stderr="")
if "fetch" in sub:
return MagicMock(returncode=0, stdout="", stderr="")
if "rev-parse" in sub and "--verify" in sub:
return MagicMock(returncode=0, stdout="abc\n", stderr="")
if cmd[:1] == ["git"] and "diff" in cmd:
diff = (
"diff --git a/src/api.py b/src/api.py\n"
"index 1..2 100644\n"
"--- a/src/api.py\n"
"+++ b/src/api.py\n"
"@@ -12 +12,1 @@\n"
"+x = parse(user_input)\n"
)
return MagicMock(returncode=0, stdout=diff, stderr="")
if Path(cmd[0]).name == "bandit":
out = json.dumps({"results": [{
"filename": "src/api.py", "line_number": 12, "line_range": [12, 12],
"issue_severity": "HIGH", "issue_text": "Use of dynamic parsing", "test_id": "B307",
}]})
return MagicMock(returncode=0, stdout=out, stderr="")
return MagicMock(returncode=0, stdout="", stderr="")
def test_cli_aborts_when_no_supported_files(tmp_path):
repo = tmp_path / "repo"
repo.mkdir()
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
return MagicMock(returncode=0,
stdout="diff --git a/README.md b/README.md\nindex 1..2 100644\n--- a/README.md\n+++ b/README.md\n@@ -1 +1,1 @@\n+x\n",
stderr="")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
with patch("shutil.which", return_value=None):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 1
manifest = json.loads((out_dir / "manifest.json").read_text())
assert any("no supported source files" in e.lower() for e in manifest["errors"])
def test_cli_happy_path_python_only(tmp_path):
from scripts import runner
repo = tmp_path / "repo"
repo.mkdir()
(repo / "src").mkdir()
(repo / "src" / "api.py").write_text("x = 1\n")
out_dir = tmp_path / "out"
mod = _load_cli()
empty_venv = tmp_path / "empty-venv"
empty_venv.mkdir()
def _which(binary):
return "/usr/bin/bandit" if binary == "bandit" else None
with patch.object(runner, "_SKILL_VENV_BIN", empty_venv):
with patch("subprocess.run", side_effect=_fake_subprocess):
with patch("shutil.which", side_effect=_which):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0
manifest = json.loads((out_dir / "manifest.json").read_text())
assert manifest["mode"] == "local"
assert manifest["language_breakdown"]["python"] == 1
rule_ids = {f["rule_id"] for f in manifest["findings"]}
assert "B307" in rule_ids
for agent in ("security-triage", "type-safety", "dependency", "consistency", "secrets"):
assert (out_dir / f"manifest-{agent}.json").exists()
def test_cli_records_tool_unavailable(tmp_path):
from scripts import runner
repo = tmp_path / "repo"
repo.mkdir()
(repo / "src").mkdir()
(repo / "src" / "api.py").write_text("x = 1\n")
out_dir = tmp_path / "out"
mod = _load_cli()
empty_venv = tmp_path / "empty-venv"
empty_venv.mkdir()
with patch.object(runner, "_SKILL_VENV_BIN", empty_venv):
with patch("subprocess.run", side_effect=_fake_subprocess):
with patch("shutil.which", return_value=None):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0
manifest = json.loads((out_dir / "manifest.json").read_text())
assert "bandit" in manifest["tools_unavailable"]
def test_cli_alerts_on_unsupported_languages(tmp_path):
"""Mixed PR: Python supported, Go file alerts — but review still proceeds."""
repo = tmp_path / "repo"
repo.mkdir()
(repo / "src").mkdir()
(repo / "src" / "api.py").write_text("x = 1\n")
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
diff = (
"diff --git a/src/api.py b/src/api.py\n"
"index 1..2 100644\n--- a/src/api.py\n+++ b/src/api.py\n"
"@@ -1 +1,1 @@\n+x = 1\n"
"diff --git a/cmd/server.go b/cmd/server.go\n"
"index 3..4 100644\n--- a/cmd/server.go\n+++ b/cmd/server.go\n"
"@@ -1 +1,1 @@\n+package main\n"
)
return MagicMock(returncode=0, stdout=diff, stderr="")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
with patch("shutil.which", return_value=None):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 0, "review should continue when at least one supported file present"
manifest = json.loads((out_dir / "manifest.json").read_text())
assert any("Go" in e and "server.go" in e for e in manifest["errors"]), manifest["errors"]
assert manifest["language_breakdown"]["python"] == 1
def test_cli_alerts_on_unsupported_only_and_aborts(tmp_path):
"""Diff contains ONLY unsupported languages: alert + abort."""
repo = tmp_path / "repo"
repo.mkdir()
out_dir = tmp_path / "out"
mod = _load_cli()
def _fake(cmd, **kw):
if cmd[:1] == ["git"] and "diff" in cmd:
diff = (
"diff --git a/cmd/server.go b/cmd/server.go\n"
"index 3..4 100644\n--- a/cmd/server.go\n+++ b/cmd/server.go\n"
"@@ -1 +1,1 @@\n+package main\n"
)
return MagicMock(returncode=0, stdout=diff, stderr="")
return _fake_subprocess(cmd, **kw)
with patch("subprocess.run", side_effect=_fake):
with patch("shutil.which", return_value=None):
rc = mod.main([
"--repo", str(repo),
"--head", "HEAD",
"--output-dir", str(out_dir),
"--mode", "local",
])
assert rc == 1
manifest = json.loads((out_dir / "manifest.json").read_text())
errors_blob = " | ".join(manifest["errors"])
assert "Go" in errors_blob, errors_blob
assert "no supported source files" in errors_blob.lower()
@@ -0,0 +1,70 @@
from scripts.diff_filter import filter_findings_by_diff
from scripts.manifest import Finding, ChangedFile
def _f(file: str, line: int, end: int = None) -> Finding:
return Finding(
tool="t", rule_id="R1", severity="medium",
file=file, line=line, end_line=end or line,
message="x",
)
def test_keeps_finding_inside_added_range():
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("src/a.py", 12)]
out = filter_findings_by_diff(findings, changed)
assert len(out) == 1
def test_drops_finding_outside_added_range():
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("src/a.py", 20)]
out = filter_findings_by_diff(findings, changed)
assert out == []
def test_drops_finding_in_unchanged_file():
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("src/b.py", 12)]
out = filter_findings_by_diff(findings, changed)
assert out == []
def test_keeps_finding_spanning_added_range():
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("src/a.py", 8, end=12)]
out = filter_findings_by_diff(findings, changed)
assert len(out) == 1
def test_keeps_finding_with_multiple_ranges():
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15), (20, 25)])]
findings = [_f("src/a.py", 22)]
out = filter_findings_by_diff(findings, changed)
assert len(out) == 1
def test_normalizes_leading_dot_slash_prefix():
"""Bandit/ruff/opengrep often emit `./src/a.py` while git diff yields `src/a.py`."""
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("./src/a.py", 12)]
out = filter_findings_by_diff(findings, changed)
assert len(out) == 1
assert out[0].file == "src/a.py"
def test_normalizes_absolute_path_with_repo_root():
"""Adapters may emit absolute paths; strip the repo root prefix."""
changed = [ChangedFile(path="src/a.py", language="python",
added_lines=[(10, 15)])]
findings = [_f("/repo/src/a.py", 12)]
out = filter_findings_by_diff(findings, changed, repo_root="/repo")
assert len(out) == 1
assert out[0].file == "src/a.py"
@@ -0,0 +1,85 @@
from unittest.mock import patch, MagicMock
from scripts.git_diff import (
resolve_default_branch, resolve_base_ref, changed_files_with_ranges,
)
_DIFF = """\
diff --git a/src/api.py b/src/api.py
index 1..2 100644
--- a/src/api.py
+++ b/src/api.py
@@ -12 +12,2 @@
-old line
+new line one
+new line two
diff --git a/src/utils/helpers.ts b/src/utils/helpers.ts
index 3..4 100644
--- a/src/utils/helpers.ts
+++ b/src/utils/helpers.ts
@@ -5,0 +6,1 @@
+ added line
"""
def test_resolve_default_branch_uses_symbolic_ref():
fake = MagicMock(returncode=0, stdout="refs/remotes/origin/main\n", stderr="")
with patch("subprocess.run", return_value=fake):
assert resolve_default_branch("/repo") == "main"
def test_resolve_default_branch_falls_back_to_master():
def _fake(cmd, **kw):
if "symbolic-ref" in cmd:
return MagicMock(returncode=1, stdout="", stderr="")
if "rev-parse" in cmd and "origin/main" in cmd:
return MagicMock(returncode=1, stdout="", stderr="")
if "rev-parse" in cmd and "origin/master" in cmd:
return MagicMock(returncode=0, stdout="abc\n", stderr="")
return MagicMock(returncode=1, stdout="", stderr="")
with patch("subprocess.run", side_effect=_fake):
assert resolve_default_branch("/repo") == "master"
def test_resolve_base_ref_uses_origin_when_available():
def _fake(cmd, **kw):
if "fetch" in cmd:
return MagicMock(returncode=0, stdout="", stderr="")
if "rev-parse" in cmd and "--verify" in cmd:
return MagicMock(returncode=0, stdout="def\n", stderr="")
return MagicMock(returncode=0, stdout="", stderr="")
with patch("subprocess.run", side_effect=_fake):
assert resolve_base_ref("/repo", "main") == "origin/main"
def test_resolve_base_ref_falls_back_when_origin_missing():
def _fake(cmd, **kw):
if "fetch" in cmd:
return MagicMock(returncode=1, stdout="", stderr="no remote\n")
if "rev-parse" in cmd and "--verify" in cmd:
return MagicMock(returncode=1, stdout="", stderr="")
return MagicMock(returncode=0, stdout="", stderr="")
with patch("subprocess.run", side_effect=_fake):
assert resolve_base_ref("/repo", "main") == "main"
def test_changed_files_with_ranges_parses_unified0_diff():
fake = MagicMock(returncode=0, stdout=_DIFF, stderr="")
with patch("subprocess.run", return_value=fake):
out = changed_files_with_ranges("/repo", "origin/main", "HEAD")
paths = {p: ranges for p, ranges in out}
assert "src/api.py" in paths
assert paths["src/api.py"] == [(12, 13)]
assert "src/utils/helpers.ts" in paths
assert paths["src/utils/helpers.ts"] == [(6, 6)]
def test_changed_files_passes_no_ext_diff_flag():
fake = MagicMock(returncode=0, stdout="", stderr="")
with patch("subprocess.run", return_value=fake) as p:
changed_files_with_ranges("/repo", "origin/main", "HEAD")
args = p.call_args[0][0]
assert "--no-ext-diff" in args
assert "--unified=0" in args
assert "diff.noprefix=false" in " ".join(args)
@@ -0,0 +1,135 @@
from scripts.language_detect import (
detect_language, is_dep_manifest, is_gha_file, is_supported_source,
known_unsupported_language,
)
def test_known_unsupported_languages():
assert known_unsupported_language("cmd/main.go") == "Go"
assert known_unsupported_language("lib/foo.rb") == "Ruby"
assert known_unsupported_language("src/Main.java") == "Java"
assert known_unsupported_language("src/main.rs") == "Rust"
assert known_unsupported_language("src/App.kt") == "Kotlin"
def test_known_unsupported_returns_none_for_supported():
assert known_unsupported_language("src/api.py") is None
assert known_unsupported_language("src/index.ts") is None
def test_known_unsupported_returns_none_for_non_source():
assert known_unsupported_language("README.md") is None
assert known_unsupported_language("Dockerfile") is None
assert known_unsupported_language("config.yaml") is None
def test_python_file_detected():
assert detect_language("src/api.py") == "python"
def test_typescript_extensions():
assert detect_language("src/index.ts") == "typescript"
assert detect_language("src/component.tsx") == "typescript"
def test_javascript_extensions():
assert detect_language("src/index.js") == "javascript"
assert detect_language("src/component.jsx") == "javascript"
assert detect_language("src/module.mjs") == "javascript"
assert detect_language("src/legacy.cjs") == "javascript"
def test_csharp_file_detected():
assert detect_language("Program.cs") == "csharp"
def test_unsupported_source_returns_none():
assert detect_language("main.go") is None
assert detect_language("script.rb") is None
assert detect_language("Main.java") is None
def test_non_source_files_return_none():
assert detect_language("README.md") is None
assert detect_language("Dockerfile") is None
assert detect_language("config.yaml") is None
def test_is_dep_manifest_python():
assert is_dep_manifest("requirements.txt")
assert is_dep_manifest("requirements-dev.txt")
assert is_dep_manifest("pyproject.toml")
assert is_dep_manifest("poetry.lock")
assert is_dep_manifest("Pipfile.lock")
def test_is_dep_manifest_javascript():
assert is_dep_manifest("package.json")
assert is_dep_manifest("package-lock.json")
assert is_dep_manifest("yarn.lock")
assert is_dep_manifest("pnpm-lock.yaml")
def test_is_dep_manifest_csharp():
assert is_dep_manifest("MyApp.csproj")
assert is_dep_manifest("MyApp.sln")
assert is_dep_manifest("packages.config")
def test_is_supported_source():
assert is_supported_source("src/api.py")
assert is_supported_source("src/index.ts")
assert not is_supported_source("main.go")
assert not is_supported_source("README.md")
def test_lua_file_detected():
assert detect_language("src/game.lua") == "lua"
def test_lua_is_supported_source():
assert is_supported_source("src/game.lua")
def test_lua_not_in_known_unsupported():
assert known_unsupported_language("src/game.lua") is None
def test_powershell_extensions_detected():
assert detect_language("scripts/Deploy.ps1") == "powershell"
assert detect_language("module/Helpers.psm1") == "powershell"
assert detect_language("module/Helpers.psd1") == "powershell"
def test_powershell_is_supported_source():
assert is_supported_source("scripts/Deploy.ps1")
def test_powershell_not_in_known_unsupported():
assert known_unsupported_language("scripts/Deploy.ps1") is None
def test_gha_workflow_yml_detected():
assert is_gha_file(".github/workflows/ci.yml")
def test_gha_workflow_yaml_detected():
assert is_gha_file(".github/workflows/deploy.yaml")
def test_gha_actions_action_yml_detected():
assert is_gha_file(".github/actions/my-action/action.yml")
def test_non_gha_yaml_not_detected():
assert not is_gha_file("config/app.yml")
assert not is_gha_file("docker-compose.yaml")
assert not is_gha_file("README.md")
def test_gha_file_wrong_extension():
assert not is_gha_file(".github/workflows/ci.json")
def test_gha_file_windows_style_path():
assert is_gha_file(".github\\workflows\\ci.yml")
@@ -0,0 +1,138 @@
"""Tests for scripts/log-run.py."""
from __future__ import annotations
import importlib.util
import json
from pathlib import Path
_SPEC = importlib.util.spec_from_file_location(
"log_run",
Path(__file__).parent.parent / "scripts" / "log-run.py",
)
log_run = importlib.util.module_from_spec(_SPEC)
_SPEC.loader.exec_module(log_run)
def _write_usage(tmp_path: Path, usage: dict) -> Path:
p = tmp_path / "usage.json"
p.write_text(json.dumps(usage), encoding="utf-8")
return p
def _read_log(log: Path) -> list[dict]:
return [json.loads(line) for line in log.read_text(encoding="utf-8").splitlines() if line.strip()]
def test_writes_one_row_per_agent_with_finding_counts(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
(output / "findings-security-triage-reviewer.json").write_text(
json.dumps({"findings": [{"file": "a.py", "line": 1, "rule_id": "B1"},
{"file": "b.py", "line": 2, "rule_id": "B2"}]}),
encoding="utf-8",
)
(output / "findings-walkthrough-reviewer.json").write_text(
json.dumps({"overview": "summary", "files": []}),
encoding="utf-8",
)
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"security-triage-reviewer": {"model": "sonnet", "input_tokens": 100,
"output_tokens": 50, "duration_ms": 1234},
"walkthrough-reviewer": {"model": "sonnet", "input_tokens": 200,
"output_tokens": 80, "duration_ms": 4321},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "abc123",
"--repo", "/tmp/repo", "--mode", "local",
"--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
rows = _read_log(log)
assert len(rows) == 2
by_agent = {r["agent"]: r for r in rows}
assert by_agent["security-triage-reviewer"]["finding_count"] == 2
assert by_agent["security-triage-reviewer"]["input_tokens"] == 100
assert by_agent["security-triage-reviewer"]["model"] == "sonnet"
assert by_agent["walkthrough-reviewer"]["finding_count"] == 0
assert by_agent["walkthrough-reviewer"]["duration_ms"] == 4321
def test_missing_findings_file_counts_zero(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"phantom-reviewer": {"model": "sonnet", "input_tokens": 0,
"output_tokens": 0, "duration_ms": 0},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "ref", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
rows = _read_log(log)
assert rows[0]["finding_count"] == 0
def test_malformed_findings_file_counts_zero(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
(output / "findings-broken-reviewer.json").write_text("{not json", encoding="utf-8")
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"broken-reviewer": {"model": "haiku", "input_tokens": 10,
"output_tokens": 5, "duration_ms": 100},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "local", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
rows = _read_log(log)
assert rows[0]["finding_count"] == 0
def test_findings_as_bare_list_is_counted(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
(output / "findings-x-reviewer.json").write_text(
json.dumps([{"rule_id": "A"}, {"rule_id": "B"}, {"rule_id": "C"}]),
encoding="utf-8",
)
log = tmp_path / "runs.jsonl"
usage = _write_usage(tmp_path, {
"x-reviewer": {"model": "sonnet", "input_tokens": 1,
"output_tokens": 1, "duration_ms": 1},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "local", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
rows = _read_log(log)
assert rows[0]["finding_count"] == 3
def test_log_path_parent_is_created(tmp_path: Path) -> None:
output = tmp_path / "out"
output.mkdir()
log = tmp_path / "nested" / "deeper" / "runs.jsonl"
usage = _write_usage(tmp_path, {
"x-reviewer": {"model": "sonnet", "input_tokens": 0,
"output_tokens": 0, "duration_ms": 0},
})
rc = log_run.main([
"--output-dir", str(output), "--run-id", "x", "--repo", "/r",
"--mode", "local", "--log-path", str(log), "--usage-json", str(usage),
])
assert rc == 0
assert log.exists()
@@ -0,0 +1,80 @@
import json
from scripts.manifest import (
Manifest, Finding, ChangedFile, LanguageBreakdown,
PackageDiff, PackageEntry, PackageUpgrade, ToolStat,
)
def test_finding_to_dict_includes_all_fields():
f = Finding(
tool="bandit", rule_id="B608", severity="high",
file="src/api.py", line=45, end_line=45,
message="Possible SQL injection.",
cwe="CWE-89", fix_suggestion=None,
)
d = f.to_dict()
assert d["tool"] == "bandit"
assert d["rule_id"] == "B608"
assert d["cwe"] == "CWE-89"
assert "fix_suggestion" in d
def test_changed_file_serializes_ranges():
cf = ChangedFile(
path="src/api.py", language="python",
added_lines=[(12, 14), (45, 45)],
)
d = cf.to_dict()
assert d["language"] == "python"
assert d["added_lines"] == [[12, 14], [45, 45]]
def test_manifest_roundtrip():
m = Manifest(
mode="local", base_ref="origin/main", head_ref="HEAD",
default_branch="main",
language_breakdown=LanguageBreakdown(
python=2, javascript=0, typescript=0, csharp=0,
skipped_files=[],
),
changed_files=[],
findings=[],
package_diffs={
"python": PackageDiff(),
"javascript": PackageDiff(),
"csharp": PackageDiff(),
},
tool_stats={},
tools_unavailable=[],
errors=[],
)
payload = json.loads(m.to_json())
assert payload["mode"] == "local"
assert payload["language_breakdown"]["python"] == 2
assert payload["findings"] == []
def test_package_diff_with_entries():
pd = PackageDiff(
added=[PackageEntry(name="pyyaml", version="6.0")],
removed=[],
upgraded=[PackageUpgrade(name="requests", from_version="2.28.0", to_version="2.31.0")],
)
d = pd.to_dict()
assert d["added"][0]["name"] == "pyyaml"
assert d["upgraded"][0]["from"] == "2.28.0"
assert d["upgraded"][0]["to"] == "2.31.0"
def test_tool_stat_records_ran_and_counts():
ts = ToolStat(ran=True, pre_filter=47, post_filter=3)
d = ts.to_dict()
assert d == {"ran": True, "pre_filter": 47, "post_filter": 3}
def test_tool_stat_records_skip_reason():
ts = ToolStat(ran=False, reason="not on PATH")
d = ts.to_dict()
assert d["ran"] is False
assert d["reason"] == "not on PATH"
@@ -0,0 +1,75 @@
from scripts.slicing import slice_for_agent
def _manifest():
return {
"mode": "local",
"base_ref": "origin/main",
"head_ref": "HEAD",
"default_branch": "main",
"language_breakdown": {"python": 1, "javascript": 0, "typescript": 0, "csharp": 0, "skipped_files": []},
"changed_files": [{"path": "foo.py", "language": "python", "added_lines": [[10, 50]]}],
"findings": [
{"tool": "vulture", "rule_id": "vulture:90pct", "severity": "medium",
"file": "foo.py", "line": 10, "end_line": 10, "message": "unused function"},
{"tool": "radon", "rule_id": "radon:cc=12", "severity": "medium",
"file": "foo.py", "line": 20, "end_line": 40, "message": "high CCN"},
{"tool": "interrogate", "rule_id": "interrogate:missing-docstring", "severity": "low",
"file": "foo.py", "line": 5, "end_line": 5, "message": "missing docstring"},
{"tool": "lizard", "rule_id": "lizard:ccn=22", "severity": "high",
"file": "foo.py", "line": 30, "end_line": 60, "message": "high CCN"},
{"tool": "knip", "rule_id": "knip:dead-export", "severity": "medium",
"file": "src/a.ts", "line": 12, "end_line": 12, "message": "unused export"},
{"tool": "jscpd", "rule_id": "jscpd:clone-40lines", "severity": "medium",
"file": "foo.py", "line": 100, "end_line": 140, "message": "duplicate"},
{"tool": "bandit", "rule_id": "B608", "severity": "high",
"file": "foo.py", "line": 5, "end_line": 5, "message": "SQL injection"},
{"tool": "ruff-idiom", "rule_id": "C901", "severity": "medium",
"file": "foo.py", "line": 50, "end_line": 50, "message": "too complex"},
{"tool": "ruff-idiom", "rule_id": "PLR0915", "severity": "medium",
"file": "foo.py", "line": 60, "end_line": 60, "message": "too many statements"},
{"tool": "ruff-idiom", "rule_id": "SIM117", "severity": "low",
"file": "foo.py", "line": 70, "end_line": 70, "message": "combine with"},
{"tool": "ruff-idiom", "rule_id": "PLR0913", "severity": "medium",
"file": "foo.py", "line": 80, "end_line": 80, "message": "too many args"},
],
"package_diffs": {"python": {"added": [], "removed": [], "upgraded": []}},
"tool_stats": {},
"tools_unavailable": [],
"errors": [],
}
def test_maintainability_slice_includes_dedicated_tools():
sliced = slice_for_agent(_manifest(), "maintainability")
tools = {f["tool"] for f in sliced["findings"]}
assert tools >= {"vulture", "radon", "interrogate", "lizard", "knip", "jscpd"}
def test_maintainability_slice_includes_complexity_idioms():
sliced = slice_for_agent(_manifest(), "maintainability")
rules = {f["rule_id"] for f in sliced["findings"]}
assert "C901" in rules
assert "PLR0915" in rules
def test_maintainability_slice_excludes_security_and_dependency():
sliced = slice_for_agent(_manifest(), "maintainability")
tools = {f["tool"] for f in sliced["findings"]}
assert "bandit" not in tools
def test_maintainability_slice_excludes_non_complexity_idioms():
sliced = slice_for_agent(_manifest(), "maintainability")
rules = {f["rule_id"] for f in sliced["findings"]}
assert "SIM117" not in rules
assert "PLR0913" not in rules
def test_consistency_slice_excludes_complexity_idioms():
sliced = slice_for_agent(_manifest(), "consistency")
rules = {f["rule_id"] for f in sliced["findings"]}
assert "C901" not in rules
assert "PLR0915" not in rules
assert "SIM117" in rules
assert "PLR0913" in rules
@@ -0,0 +1,48 @@
from scripts.package_diff import diff_requirements_txt, diff_package_json
def test_requirements_txt_diff():
before = "requests==2.28.0\npyyaml==5.4.1\n"
after = "requests==2.31.0\npyyaml==5.4.1\nclick==8.0.0\n"
pd = diff_requirements_txt(before, after)
assert {p.name for p in pd.added} == {"click"}
assert {p.name for p in pd.removed} == set()
assert len(pd.upgraded) == 1
assert pd.upgraded[0].name == "requests"
assert pd.upgraded[0].from_version == "2.28.0"
assert pd.upgraded[0].to_version == "2.31.0"
def test_requirements_txt_removed():
before = "a==1.0.0\nb==1.0.0\n"
after = "a==1.0.0\n"
pd = diff_requirements_txt(before, after)
assert {p.name for p in pd.removed} == {"b"}
assert pd.added == [] and pd.upgraded == []
def test_package_json_diff():
before = '{"dependencies": {"axios": "0.21.0", "lodash": "4.17.20"}}'
after = '{"dependencies": {"axios": "1.6.0", "react": "18.0.0"}}'
pd = diff_package_json(before, after)
assert {p.name for p in pd.added} == {"react"}
assert {p.name for p in pd.removed} == {"lodash"}
assert pd.upgraded[0].name == "axios"
def test_package_json_handles_dev_dependencies():
before = '{"dependencies": {}, "devDependencies": {"jest": "27.0.0"}}'
after = '{"dependencies": {}, "devDependencies": {"jest": "29.0.0"}}'
pd = diff_package_json(before, after)
assert pd.upgraded[0].name == "jest"
assert pd.upgraded[0].to_version == "29.0.0"
def test_empty_before_treats_all_as_added():
pd = diff_requirements_txt("", "click==8.0.0\n")
assert {p.name for p in pd.added} == {"click"}
def test_invalid_json_returns_empty_diff():
pd = diff_package_json("not json", "not json")
assert pd.added == [] and pd.removed == [] and pd.upgraded == []
@@ -0,0 +1,52 @@
import json
from pathlib import Path
from scripts.review_stats import compute_stats
def _write_log(path: Path, records: list[dict]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w") as f:
for r in records:
f.write(json.dumps(r) + "\n")
def test_precision_per_agent(tmp_path):
log = tmp_path / "runs.jsonl"
_write_log(log, [
{"kind": "subagent_run", "run_id": "r1", "agent": "sec",
"model": "claude-sonnet-4-6",
"input_tokens": 1000, "output_tokens": 200,
"duration_ms": 100, "finding_count": 4, "ts": "x", "repo": "/r", "mode": "local"},
{"kind": "verdict", "run_id": "r1", "agent": "sec", "rule_id": "B1", "file": "a", "line": 1, "verdict": "kept"},
{"kind": "verdict", "run_id": "r1", "agent": "sec", "rule_id": "B2", "file": "a", "line": 2, "verdict": "kept"},
{"kind": "verdict", "run_id": "r1", "agent": "sec", "rule_id": "B3", "file": "a", "line": 3, "verdict": "dismissed"},
{"kind": "verdict", "run_id": "r1", "agent": "sec", "rule_id": "B4", "file": "a", "line": 4, "verdict": "false_positive"},
])
stats = compute_stats(log)
sec = stats["by_agent"]["sec"]
assert sec["kept"] == 2
assert sec["total"] == 4
assert abs(sec["precision"] - 0.5) < 1e-9
assert sec["tokens"] == 1200
assert sec["tokens_per_kept"] == 600.0
def test_per_rule_precision(tmp_path):
log = tmp_path / "runs.jsonl"
_write_log(log, [
{"kind": "verdict", "run_id": "r1", "agent": "sec", "rule_id": "B101", "file": "a", "line": 1, "verdict": "dismissed"},
{"kind": "verdict", "run_id": "r2", "agent": "sec", "rule_id": "B101", "file": "b", "line": 1, "verdict": "dismissed"},
{"kind": "verdict", "run_id": "r3", "agent": "sec", "rule_id": "B101", "file": "c", "line": 1, "verdict": "false_positive"},
])
stats = compute_stats(log)
rule = stats["by_rule"]["sec/B101"]
assert rule["kept"] == 0
assert rule["total"] == 3
assert rule["precision"] == 0.0
def test_empty_log_returns_zero_stats(tmp_path):
log = tmp_path / "runs.jsonl"
stats = compute_stats(log)
assert stats == {"by_agent": {}, "by_rule": {}, "runs": 0}

Some files were not shown because too many files have changed in this diff Show More