team · models

Which model looks for problems, and which one judges them

«Which model should I use for code review» is the wrong question, and the price tag shows why. On our own runs the judge takes about half of what a review costs — so you are not choosing a model, you are choosing models for two different roles. One looks at the diff and proposes what might be wrong. The other reads whole files and decides which of those proposals survive.

The roles want opposite things. The generator needs breadth and a low price: a model that proposes ten findings, three of them noise, has done its job. The judge needs stubbornness: a model that rejects everything it cannot prove from the code. Give both roles to the same model, and the second pass will mostly agree with the first — exactly the behaviour you wanted to avoid.

Before you start: this is about the llm block of .reviewgate/config.yml — the config reference lists every key, this page is about which of them to fill and why. Which schema is in effect for you right now, and where it came from, reviewgate doctor will show.

The roles in a review schema

RoleHow manyWhat it seesWhat it decides
generatorsone or morethe diff, plus whole files if you enabled thatwhat might be wrong — it proposes, nothing more
judgesnone, one or twowhole files of every finding, the files they import, and the exact lines a ready-made fix would replacewhich proposals survive, and at what severity
arbiternone or oneonly the findings the two judges disagreed aboutwhich judge is right

Without a config you get the minimum of this table: one generator, no judges. Everything it finds comes to you unverified — fine for a draft on your own machine, thin for a merge or pull request.

What a second generator buys

Two generators do not vote. Their findings are merged and deduplicated, so the second one adds what the first missed rather than confirming what it found. Different vendors are worth more here than a bigger model from the same one: models trained differently miss different things, and that is the entire point of the pair. A second reason: code written by an AI agent should not be checked by the same model — it would repeat its own mistakes. Two generators from different vendors make sure at least one of them looks at the diff with different eyes.

What it costs is another full call on the same input — and the surprise is how little that is in money and how much in time. We measured it — the table is below.

What a second judge buys

Two judges do vote. Where they agree, the verdict stands. Where they disagree, the finding goes to the arbiter — and if there is no arbiter, the disputed finding is dropped conservatively. On earlier measurements the splits were real, not theoretical: in three runs out of five the judges disagreed — on four findings in one run, and on one and on two in the other two.

two configurations that do nothingAn arbiter with no judges is removed from the run — there is nothing to arbitrate. An arbiter with exactly one judge stays in the config and is never called: one judge cannot disagree with itself. Both cases produce a notice rather than silence, because a dead key that reads like a working setting is worse than an error.

What we measured, and what it does not prove

Eighteen runs, on 7 and 22 September 2026: six schemas, three runs each, the same 31-line diff. We picked a diff where the right answer is known in advance — four real defects: fractional cents, a customer's e-mail in a log line, tax computed before the discount, and a coupon lookup on a plain object that also reaches inherited prototype properties. In the four bottom rows a second generator is added to the pair «sonnet-5 + judge opus-5» — that is what we compare.

SchemaFindingsOf 4 defectsCostTime
one generator, no judge3.02.7$0.1067 s
generator + judge3.73.0$0.21109 s
+ DeepSeek flash3.03.0$0.23210 s
+ opus-5 (bigger, same vendor)4.33.3$0.39154 s
+ gpt-5.6-terra (other vendor, same tier)4.33.3$0.26128 s
+ gpt-5.6-sol (other vendor, bigger)5.34.0$0.27198 s

The judge adds no coverage — it can only drop and downgrade, it cannot find anything. The gap between 2.7 without a judge and 3.0 with one is generator variance between runs: the fourth defect was found by one run out of three, and it was the generator that found it. That is not a disappointment, it is the division of labour: the judge works on precision, not on search. In the three «generator + judge» runs it dropped none of the 11 candidates and lowered a severity in two; it starts dropping when there are two generators — the second one's duplicates and guesses.

A second generator from another vendor adds coverage — if it is no weaker than the first. The weak DeepSeek flash added nothing: its two or three findings per run either duplicated sonnet-5's or were dropped by the judge. gpt-5.6-terra — the same tier as sonnet-5 — gave exactly as much as the bigger model of the same vendor, opus-5: 4.3 findings and 3.3 known defects, but 33% cheaper and faster. The bigger model of the other vendor, gpt-5.6-sol, found all four defects in each of the three runs — 5.3 findings against opus-5's 4.3, and 30% cheaper.

A second generator costs cents; what you pay is time. Terra and sol added 5–6 cents to a run, opus-5 — 18; in seconds — 19, 89 and 45: the second generator runs after the first, not in parallel with it, and its time is added in full. Time varies more than price: sol answered in 38 to 152 seconds, DeepSeek in 51 to 161.

What this does not prove: three runs per schema, one diff, one language — this is a measurement, not a study. Read it as «this is what happened on our repository», not as a model ranking. The judge is the same in every schema, opus-5 — in the schema where it is also the second generator it judged its own candidates. Anthropic and DeepSeek prices are per their price lists on 23 September 2026 (DeepSeek at the peak rate), the GPT-5.6 prices per OpenRouter's price list on the day of the measurement; check your provider's price list.

Three places a schema can live, and they do not merge the way you expect

A schema can come from the team config (.reviewgate/config.yml), from your personal config, or from the installation's environment. The rules differ per source, and that difference is where people go wrong.

Between the team config and the environment the roles resolve independently. A run can legitimately use generators from config.yml, judges from the environment and an arbiter from there too. Nobody wrote that combination down anywhere; it is assembled per role.

Your personal schema is taken whole and overrides the team's. Declare even one role in your personal config — say, only generators — and the whole trio comes from there, including the roles you did not write. No judges in your personal file means no judges at all, not the team's panel. Your local run then looks cheaper and cleaner than the review on the merge or pull request, and the difference is not a bug: you changed the schema, and the repository stayed the same.

This is deliberate — a personal schema that silently half-merges with the team's would be worse — but it means one thing for you: a local run with a personal schema does not predict what the bot will say. The run says so in a notice, and the notice is easy to miss.

terminal
reviewgate review --refs main --team-llm

--team-llm drops your personal schema for one run. Removing the role keys from your personal config drops it permanently. And a run through the team's review server always uses the team schema anyway, because the bot never reads your file.

A role can name a vendor — backend: deepseek — but a backend is a name of an entry in the installation's catalogue, not an address and not a secret. The keys sit in the installation's environment; the config in the repository can only point at them. That is the invariant that lets you clone a stranger's repository without their config sending your diff to their endpoint.

The one rule that matters

The judge must be stronger than the generator. Make a weaker model the judge, and it will drop good findings it cannot follow while agreeing with confident nonsense — you pay for a second pass that makes the result worse. The effect is strongest when the judge is not only stronger but from another vendor: the same logic as for the second generator — it does not share the generator's blind spots. Everything else in this section is preference; this is the rule.

The second useful habit: when you add a second generator, take a different vendor rather than a bigger model from the same one — and no weaker than the first. You are buying different blind spots, and two models of one family share theirs; the table above shows it: terra gave as much as opus-5 for less money, and sol found the most.

.reviewgate/config.yml
llm:
  generators:
    - model: claude-sonnet-5          # the main pair of eyes
    - backend: openrouter             # a second vendor, no weaker than the first
      model: openai/gpt-5.6-sol
  judges:
    - model: claude-opus-5            # verifies against whole files
    - backend: openrouter             # a judge from another vendor: other blind spots
      model: openai/gpt-5.6-sol
  arbiter:
    model: claude-fable-5-1           # called only where the two judges disagree

This is an example of a full schema, not what you need from day one: two vendors in both roles plus an arbiter above them — the most expensive configuration possible. Our own repository lives on one generator and one judge. A sensible starting point is one generator and one stronger judge; the next step is a second generator from another vendor, and by the table above it pays for itself.

What each role costs

The judge is the most expensive role in the schema, and the runs above prove it: $0.21 with it against $0.10 without, so it is 53% of the bill. The other lever is scope — the judge reads whole files, so its share grows with the size of the files your findings land in, not with the size of your diff.

The local option. Two different cases here. A single model on your own hardware — such as qwen3-coder:30b on a workstation — costs nothing per run, but runs several times slower than the cloud pair and is no match for it: on the same diff it found none of the four known defects. That is an option for an enthusiast, not for a team. A company in a closed perimeter can run stronger models — open DeepSeek, Qwen and the like — and then the «generator + judge» schema works without the cloud; you pay not per run but once, for the hardware and for deploying several models, and that is serious money. Treat local models as the answer to «the code cannot leave the perimeter», not as a cheap version of the same thing.

When it goes wrong

Your local run and the bot disagree on the same branch. Almost always it is the home schema: it replaces the team's trio whole, so you and the bot are on different schemas. Check reviewgate doctor — it shows where each role came from — then either remove the role keys from the home config or run once with --team-llm.

--full runs without a judge. The --full flag does not add a judge, it only stops removing one. If no judge is configured in the team file, the environment or your personal file, the run goes as --fast with a notice — the exit code and the verdict look exactly the same as a judged run.

The run refuses to start with home_llm_config. Your personal schema is declared but has no generators. The tool will not quietly borrow generators from elsewhere — that would assemble a configuration nobody designed. Add a generator, remove the role keys, or bypass once with --team-llm.

The arbiter is configured and never appears. With no judges it is removed from the run; with exactly one judge it is kept and never called. Both are reported as notices — read them, because from the config alone the key looks alive.

A role is ignored. A role entry with neither model nor backend addresses nothing and is dropped with a notice. Same for a backend name that is not in the installation's catalogue.

The second generator contributed nothing. Most likely it did not answer — it ran out of time or returned an error — and the ensemble collapsed to the models that did; the run's log has the line «failed — skipping it (fail-open)». The diagnostics block shows every call with its duration and outcome, and that is how you tell «the model found nothing» from «the model never answered». It is off by default: turn diagnostics on and repeat the run.

Next