Write review rules your linter cannot check
Here is a diff from our sample repository (on GitLab or on GitHub): ESLint says nothing, seventeen tests pass, TypeScript is happy — and the coupon discount is computed as (net.amount * percent) / 100, so an invoice for €9.99 quietly acquires a fraction of a cent. The review found it in a minute and a half and rated it 🔴 critical, because the team had once written down one line: «Amounts live in integer minor units…».
That line is what this page is about. Not «configure the rules section» — but which of your team's agreements are worth writing down at all, how to word them so the review acts on them, and what the review will never do with them no matter how you word it.
Before you start: you need a repository with .reviewgate/config.yml — if there is none yet, your first local review shows where it comes from. To see the rules in effect: reviewgate rules.
A rule is not a setting — it is a sentence you send to the model
A rule has exactly three fields: id, description, severity (major if you omit it). The text goes into the system prompt verbatim, one line per rule:
- [critical] money-in-minor-units: Amounts live in integer minor units…
- [major] no-customer-data-in-logs: Logs carry identifiers and outcomes…Two consequences people do not expect. The model sees the severity you wrote, so it is not merely post-processing — it colours how the finding is judged in the first place. And the same block goes to the generator, to every judge and to the arbiter: a sprawling rulebook is paid for on every call of every run, not once.
The filter: three places an agreement can live
Before writing a rule, ask what could check the agreement. Only three things can — the linter, the tests or rules[] — and only the last one does it by understanding what the code means.
| The agreement | Where it belongs | Why |
|---|---|---|
«no any outside tests», «=== not ==», «no floating promises» | linter | a machine checks this deterministically, at no cost, in milliseconds — and fixes some of it automatically |
| «a discount of 10% off 24.80 gives 2.48», «an invoice issued on the evening of the 31st, customer's time, lands in that month, not in the next one by UTC» | tests | it is behaviour, and behaviour is expressed as an assertion, not as a sentence |
| «money lives in integer minor units», «every database read is filtered by the tenant the request came from», «logs carry identifiers, never customer data» | rules[] | no tool can check it: it needs to understand what this code is doing and why |
The dividing line is not «important or unimportant». It is «can a program verify this without understanding the meaning». A rule that repeats the linter costs you money on every call and buys nothing — the linter already said it, faster and at no cost.
A fourth source is the one people duplicate without noticing: the preset. preset: nestjs or preset: angular already puts a page of instructions into the prompt — floating promises, defaults through || where falsy values are valid, framework-specific traps. The preset text itself is not exposed: the command below shows only its name, and what each preset checks is listed in the preset table of the config reference.
reviewgate rulesWrite the exception into the rule, not around it
A rule written without exceptions argues with your own architecture: the review keeps reporting a pattern your team chose deliberately, and every such report costs exactly what a real finding costs. One sentence inside the rule text fixes it, because the exception now lives where the model reads it. The same trick replaces most of what you would otherwise put into dont_flag.
- id: no-direct-db-access
description: "Services do not touch the database directly"- id: no-direct-db-access
description: "Services go through a repository. Exceptions: migrations and seed scripts
under db/, where direct access is intentional."Say what counts as proof
A rule about something the code does not show — a missing file, an absent test, a convention about the diff as a whole — is something the model cannot verify by reading the code. If the rule does not say what counts as proof, the model decides for itself: in our sample repository the first rule below produced a finding that asks for a link to a finance ticket and a sign-off from finance — nobody asked for either, and nobody can check them. The same rule with the evidence named produces a finding about exactly what is missing: the absence of README.md among the changed files is the evidence, not a hypothesis.
- id: pricing-changes-documented
description: "Every change to pricing rules (discounts, tax rates) must be documented
for the finance team."🟠 src/pricing/pricing.service.ts:18 — A new tax rate (CH: 0.081) is added to TAX_RATES,
which is a pricing rule change affecting tax computed on every Swiss invoice. The team
rule requires every pricing change (discounts, tax rates) to be documented for the
finance team, but no documentation, changelog entry, or reference to a finance sign-off
accompanies this change. (team:pricing-changes-documented)
Add a comment or link to the finance ticket/documentation confirming the 8.1% rate and
the effective date, and ensure finance has signed off before merging.- id: pricing-changes-documented
description: "Every change to src/pricing/ must come with a change to README.md in the
same diff. README need not be read: the absence of README.md among the changed files
IS the evidence, not a hypothesis."🟠 src/pricing/pricing.service.ts:18 — This diff modifies src/pricing/pricing.service.ts
(adds the CH tax rate) but does not include a change to README.md in the same diff. Per
team policy, every change under src/pricing/ must be accompanied by a README.md update
in the same MR; the absence of README.md among changed files is itself the violation.
(team:pricing-changes-documented)
Add a README.md update in this MR documenting the new CH tax rate and the reasoning for
the tax table change.The id is a connector, not a name
When a finding's ruleId matches a rule exactly, three things happen: the finding gets your severity, it is marked in the summary as a team rule, and the model is instructed not to borrow your id for anything else. That is why the id has to match character for character: a stray space or a renamed rule costs you all three, and nothing warns you.
severity is a ceiling, not a guarantee. Your severity is applied before judging, and the judge may lower it — it is not allowed to raise it. So severity: blocker written to force a gate can arrive as minor and let the branch through. If a rule must hold the gate, its text has to be strong enough that a judge reading the whole file agrees with the level, not just the label.
Length costs money on every call. Nothing limits how many rules you write or how long each one is. The bill does: the block travels to the generator, to every judge and to the arbiter. Ten sharp sentences beat forty vague ones twice over — they cost less and they leave the model less room to interpret.
What rules cannot do
A rule has no scope. Rules in rules[] take no paths: a rule applies to the whole diff or not at all. If it concerns only part of the code, write the boundary into the text itself: «in the api/ layer…», «for files under legacy/ this does not apply». The clause works: the model sees the path of every file in the diff.
One preset cannot cover two stacks. A repository gets one preset, and presets cannot be set per directory. In a monorepo, pick the preset for the stack where a mistake costs more, and describe the second stack's specifics in review_prompt (mode: extend adds your text to the default guidance; replace puts it in place of the default).
Four ways to say «do not flag this»
| Tool | Where it applies | Saves tokens | Path-aware |
|---|---|---|---|
| ignore | before the model is called — the file never enters the prompt | yes | yes, globs |
| reviewgate-ignore-file | same, but decided by the file itself | yes | per file |
| dont_flag | a sentence in the prompt asking the model not to raise a class of findings | no | no |
| min_severity | after the answer: quiet findings fold into the summary | no | no |
The row people misread is dont_flag. It is not a filter: nothing is removed after the answer. It is text in the prompt: the generator is told not to raise findings these assumptions cover, and the judge to drop the ones that contradict them. It works well, and it costs exactly what any other sentence in the prompt costs. If your goal is money, only the top two rows help.
The row that surprises people is min_severity. It looks like a noise filter for the feed, but muted findings leave the gate too: with min_severity: major, a run with five 🔵 minor findings stops neither a commit nor a merge request, whatever gate threshold you set. And the two surfaces count them differently: the CLI verdict line says 0 while still listing all five marked «muted», and the bot's summary on a merge or pull request counts all five and folds them into a collapsed block. One policy, two numbers.
tests: optional tells the model not to flag missing tests as a problem, so the team is not nagged about coverage. A team rule is stronger: with the flag on, a rule like «a new service comes with a test» still fires — in our sample repository the review raised the missing test and the judge kept it. If you want silence about tests, do not write such rules; if you want tests checked selectively, keep the flag and write the rule.What a team rule looks like when it fires
The same coupon diff as at the top of the page, in a repository that has the rules:
🔴 src/pricing/pricing.service.ts:52 — couponDiscount computes
(net.amount * percent) / 100 directly instead of going through the shared
multiply/roundHalfUp helper. Because minor units are integers, this division can
yield a fractional amount (net=101, percent=7 -> 7.07), silently violating the money
invariant. (team:money-in-minor-units)
return multiply(net, percent / 100);The tail in brackets is the whole point: the finding is not the model's opinion, it is your agreement, quoted back at the line that breaks it. Review conversations end faster when nobody has to argue about whether this matters — the team already decided.
Run reviewgate rules after every edit of the config, for one reason: a broken rule drops out quietly enough to miss. A rule has two required fields, id and description; if either is missing or named differently, the whole rule is skipped. There is a notice in the summary, but it is one line among many, and nothing else changes: the review runs, the report looks normal, the rule simply is not in it. The first line of the command's output shows how many rules were loaded (rules=6); if the number is smaller than the count in your file, one of them did not parse — that is the whole check. Duplicated ids are worse: both rules go to the model, but the severity of the last one wins, without a single warning.
When it goes wrong
The rule is written, but no finding ever cites it. Three possible causes, in order of likelihood: the rule was dropped on parsing — check with reviewgate rules; the id does not match what the model produced, character for character; or the rule is worded so vaguely that the model did not connect the finding to it — say in the rule text what exactly counts as a violation.
The review keeps flagging legacy code you have no intention of fixing. The rule has no exception clause, and rules have no path scope. Write the boundary into the sentence — «this does not apply under legacy/» — rather than adding another dont_flag line.
Every run warns about your config. Rules accept exactly three fields. Anything else — globs, prompt, files — produces a notice on every single run. An AI assistant editing your config is the usual source: it writes the schema it expects, not the one that exists. Path scope has no key on purpose: a rule applies to the whole diff, and the boundary goes into its text.
A rule marked blocker does not hold the gate. Either the judge lowered the severity — your level is a ceiling, not a floor — or min_severity stands above your gate threshold and mutes the finding out of the gate.
Next
- Your first local review— where the config comes from
- The hook ladder— making the gate act on these rules
- Config reference— every key, including dont_flag, min_severity and review_prompt