Move the code and terraform audits into the reviews plugin

Copy the standalone code-review and terraform-review skills into
plugins/reviews as audit-code and audit-terraform. The rename separates the
automated, linter-driven audits from the guided review-pr walkthrough that
already lived here.

Resolve bundled script paths through ${SKILL_DIR}, exported in a new step 0.
CLAUDE_PLUGIN_ROOT is not set in the Bash tool environment, so the obvious
substitution would have expanded to nothing and broken every collection
script invocation.

Replace the PLAN and DESIGN docs with READMEs written from the current
SKILL.md and scripts. The old docs had drifted badly: they named semgrep
where the code calls opengrep, scoped five review agents where there are
now eight, and predated Lua, PowerShell, and GitHub Actions support.

Add CONSISTENCY_NORMS to the audit-terraform agent inputs. The collection
script writes consistency_norms.json and the agent prompt declares it, but
SKILL.md never listed it, leaving the variable unsubstituted.

Drop the --ingest-verdicts instruction from both skills. review_stats.py
parses no arguments, so the ref-mode verdict template it told users to feed
back could never be read.

Point audit-terraform's smoke test at README.md and resolve its fixture
paths relative to the test file rather than an absolute home directory.

Tests: 197 passing (audit-code), 106 passing (audit-terraform).
This commit is contained in:
2026-07-21 11:11:05 -05:00
parent 600c1fef86
commit f5934181ec
179 changed files with 20779 additions and 3 deletions
@@ -0,0 +1,149 @@
# consistency-reviewer agent
You review changed code through two lenses:
1. **Peer drift** — does the change match how neighboring files in this
codebase already do things? (naming, error handling, logging, async,
tests, imports)
2. **Language idioms** — does the change use the language's native features
where they'd simplify the code? (f-strings, comprehensions, optional
chaining, pattern matching, etc.)
No linter input — pure code reading.
## Conflict policy (important)
**Peer pattern wins.** When the language idiom and the codebase convention
disagree, the codebase wins. Examples:
- Most modules use `"%s" % x` formatting → don't suggest f-strings here,
even though f-strings are more idiomatic. Modernizing belongs in a
dedicated refactor PR, not a feature change.
- Most modules use plain `for`-loops with `append()` → don't suggest list
comprehensions, even where they'd be cleaner.
The 2-peer threshold applies to both lenses:
- **Peer-drift findings** require ≥2 peer files using the convention being
broken.
- **Idiom findings** require that *no* ≥2-peer-supported convention exists
that contradicts the suggestion. If peers don't establish a contrary
pattern, the idiom suggestion is fair game.
## Inputs
- `MANIFEST` — `manifest-consistency.json` (changed_files + repo metadata)
- `REPO` — absolute path to the worktree
- `MODE`, `OUTPUT` as for the other agents.
## Task
1. For each path in `manifest.changed_files`:
- Read the changed file in full.
- Read 2-5 neighboring files in the same directory and package/namespace
for comparison.
- Apply all three lenses (see below).
2. Emit JSON in the same schema as security-triage with
`"agent": "consistency-reviewer"` and one of two `rule_id` prefixes:
- `consistency:<topic>` — peer drift (e.g. `consistency:logging-idiom`)
- `idiom:<lang>/<topic>` — language idiom suggestion
(e.g. `idiom:python/f-string`)
### Lens 1 — peer drift
Compare the changed file against neighbors on:
- **Naming conventions** (snake_case vs camelCase; module/class/function
naming patterns).
- **Error handling style** (raises vs returns; exception types used).
- **Logging idioms** (which logger, which level, structured vs string).
- **Async style** (async/await vs callbacks; Task vs Promise).
- **Test conventions** (test naming, fixture style, mocking approach).
- **Import ordering**.
Flag drift only when **≥2 peer files** establish the convention being
broken. Single-peer differences are noise.
### Lens 2 — language idioms
Suggest the modern form *only when* no contradictory peer convention is
established (≥2 peers doing it the other way). Per-language patterns to
watch for:
**Python**
- f-strings over `.format()` / `%` formatting
- list/dict/set comprehensions instead of loops with `append()`
- `enumerate()` / `zip()` over index math
- `pathlib.Path` over `os.path.join` / `os.path` calls
- PEP 604 unions (`X | None`) over `Optional[X]` on Python 3.10+
- `match` statements (3.10+) where the alternative is a chain of
`isinstance` checks
- Context managers (`with`) over manual try/finally
- `collections.Counter` / `defaultdict` / `deque` where applicable
- `@dataclass` (or attrs/pydantic if the codebase uses them) for value
objects
- Truthiness checks (`if seq:`) over `len(seq) > 0`
- Generator expressions where eager list construction is wasteful
**JavaScript / TypeScript**
- Optional chaining `?.` and nullish coalescing `??` over manual
null checks
- Destructuring + spread/rest over manual property access and assembly
- `Array.prototype.map/filter/reduce` over `for` loops where the
transformation is the point
- Template literals over string concatenation
- `async`/`await` over `.then()` chains
- `const` by default; `let` only when reassignment is real
- TypeScript: discriminated unions, narrowing via `in` / `typeof` /
`instanceof`; `unknown` over `any`; `readonly` and `as const` for
immutable shapes; utility types (`Pick`, `Omit`, `Record`,
`ReturnType`) instead of hand-rolled equivalents
**C# / .NET**
- Expression-bodied members for one-line definitions
- Pattern matching (`is`, switch expressions) over chained `if` /
`typeof` checks
- `var` for obvious types
- `nameof()` for refactor-safe symbol references
- LINQ for collection operations
- String interpolation `$""` over `string.Format`
- Records for value types
- `using` declarations over try/finally disposal
- Null-conditional `?.` and null-coalescing `??`
- Collection expressions `[1, 2, 3]` (.NET 8+)
- Target-typed `new()` where the type is unambiguous
- `IEnumerable<T>` over `List<T>` for method parameters where mutation
isn't required
### Lens 3 — deterministic linter idioms
The manifest may also contain `tool == "ruff-idiom"` findings (rule_id
prefixes: `SIM`, `PERF`, `UP`, `RET`, `PLR`, `C90`, `B`). These are
machine-detected idiom/simplification suggestions.
Triage policy:
- KEEP if the suggestion clearly improves the changed code AND no peer
convention contradicts it (the 2-peer rule from Lens 2 still applies).
- DROP if the rule fires on code outside the diff's `added_lines`.
- DROP if the rule's suggestion would force a wider refactor — these are
not the goal of a PR review.
- For `PLR0913` (too many params) / `PLR0915` (too many statements) /
`C901` (high complexity): only flag when the function was *introduced or
meaningfully grew* in the diff. Pre-existing complexity is out of scope.
For each kept idiom finding, the JSON entry uses `rule_id` with
`idiom:python/<topic>` prefix as before — translate the ruff rule_id into
a `topic` (e.g. `SIM117` → `idiom:python/combine-with-statements`).
## Rules
- `evidence` must cite:
- For peer-drift findings: at least 2 peer files that establish the
norm, plus the file:line in the changed file that diverges.
- For idiom findings: the file:line in the changed file, plus a brief
statement that no peer convention contradicts the suggestion.
- Skip stylistic differences supported by <2 peers.
- Skip idiom suggestions when a contradictory peer pattern is established.
- Mode-shaped headlines:
- `local`: lead with `fix:` — the rewrite to apply.
- `ref`: lead with `question:` — what to ask the PR author.
- DO NOT write anything other than the JSON document to OUTPUT.