team · verification

«No findings» — but did it look?

An empty summary is not one message. The bot has four different ways of saying «nothing», and they mean different things: not a single file reached the model; the model raised candidates and the judge dropped them all; the model looked and found nothing; or there are no new findings, but earlier ones are still open.

You can tell these cases apart by the summary itself: each has its own line. What the summary cannot do by default is explain why the result is empty: the diagnostics setting is off, and without it the bot keeps the most useful part of the answer to itself — which candidates were raised, and why each of them was dropped. Turn it on in advance: the next run of the same MR will be a different run — for instance, one that checks only the new changes.

For an MR that has already come back empty there is a way around: run the same branch through the CLI with --json. In that report, as in the MCP response, the dropped candidates and their reasons are always there, whether diagnostics is on or not. It will be a new run, but through the review server it goes through the same engine and the same models as the bot. The CLI's plain output does not print the dropped ones, and the hook passes the agent at most five findings — and none of the dropped.

Turning the diagnostics block on

One line in the repository config is enough:

.reviewgate/config.yml
# .reviewgate/config.yml
diagnostics: true

For a whole installation it is DIAGNOSTICS=true in the bot's environment. But if the repository config says diagnostics: false, that wins. The block needs no extra model call: it shows numbers the run has already collected, and costs nothing but a few lines of summary.

From the next run on, the summary carries a folded «🔬 Run diagnostics» block. Here is one from a real run on our sample repository:

in the summary
**Model context**: diff: 2 files · full files: 3 (6K chars) · environment: ✓ · from team standards (6 rules): 1 of 1 findings

**Calls**:

- generator `claude-sonnet-5` — 11.7K→1.2K tokens · 14 s · findings: 2
- validator `claude-opus-5` — 10.9K→186 tokens · 22 s · dropped: 1 · downgraded: 0

**The judge dropped / downgraded / removed a fix (✂️)**:

- ❌ `src/reports/reports.service.ts:20` `missing-input-validation` — «Speculative input-validation nitpick with no evidence of untrusted caller; style/defensive suggestion outside standards.»

The block has three parts, and each answers its own question about the emptiness. Model context — what the model actually received: how many files are in the diff, how many of them ignore hid, how many full files the generator got (with full_file_context), whether anything was left out by the budget, and how many findings matched your team rules. Calls — who ran, on which model, for how long, and how many findings each produced. If the generator shows findings: 0, the emptiness happened before judging: in such a run the judge is not called. The third part is what the judge dropped, downgraded or stripped of its fix, with a reason for every candidate, most often in the judge's own words.

The line above is the judge doing its job: a speculative nitpick with no evidence behind it, dropped. How to read such reasons is covered below, in «Reading the judge's verdicts».

Four different «nothing»

What the summary saysWhat happenedWhere to look
🤖 ReviewGate — no files left to review after the ignore filter. (or the marker variant) not a single file reached the model: they were all filtered out before the call your ignore list and reviewgate-ignore-file markers inside files; binary files and changes with no lines of code — renames, mode changes, deletions — end up here too. The breakdown by reason is only in the bot's log
➖ No new findings: 5 candidates were raised, all dropped by the reviewing model the generator raised candidates, and the judge dropped them all. This line only appears when a judge is configured the 🔬 block — each drop carries its reason
✅ No findings.the model looked and found nothingModel context in the block and the ⚠️ budget line in the summary: how much of the change the model actually saw
Findings reported earlier and still open (3) remain in the threads below. comes on top of the second or the third: the run found nothing new, but earlier findings are still open the threads themselves — and the gate, if it is on: they can keep it red

Only the third line can be called clean, and even that with a caveat: it does not check how complete the review was — see «Empty, but not clean». The first means there was no review, the second that there was one but the judge dropped everything, the fourth that you should be looking at the older threads. And if the summary did not update at all, that is a different case: the bot said nothing.

Reading the judge's verdicts

The judge can only take away: drop a candidate, lower its severity, or revoke its ready-made fix. It cannot add findings or raise a severity. An empty result after judging therefore means one thing: the judge dropped every candidate. The reasons for the drops fall, roughly, into three families. Besides the 🔬 block you can read them in the --json report: the dropped candidates sit in dropped — file, title and reason, with no line number.

About the finding: «speculative», «no evidence». This is the judge doing exactly what it is there for: the generator is wordy, and the judge filters out what your team has no reason to read.

About what the judge could see: «the calling code is not shown», «the content is not visible». The judge gets the files with findings and the definitions of what they import — nothing more. It never sees the code that calls the changed spot, so a finding whose proof lives there is dropped as speculation; no budget setting changes that. It parses imports in TypeScript and JavaScript, while in Python, Go, Java and many other stacks it sees only the files with findings — expect more drops of this kind there.

About your rules: «the rule does not say that», «outside the team's standards». If the findings you actually want keep getting dropped with this reason, the rule does not say what you think it says: it is the config that needs fixing, not the model.

It also happens that the judge did not run at all. The findings are then published without validation, and the summary says so itself, right under the verdict, no diagnostics block needed: «⚠️ judging failed entirely — findings are published without validation». Nothing is lost, but nothing was checked either; if ready-made fixes are on, they too go out with an Apply button nobody checked.

Empty, but not clean

An empty summary, or an MR with no new threads, does not yet mean that everything was checked and is clean. Here are four cases where it is not.

The run looked at only part of the change. The summary says so: Incremental review: 2 of 4 changed files checked (the rest are untouched since the last review). The other two files were reviewed earlier, and their findings stayed in their threads. The gate, if it is on, counts them: ❌ Severity gate (`major`) failed: 3 blocking findings (2 of them opened earlier — this run did not re-check their files).

Coverage was partial. Some files never reached the model, and the summary reports it in different ways. The budget gets a line of its own: «⚠️ This MR exceeds the review budget: 4 of 11 files covered». An in-file marker gets one too: «ℹ️ Excluded by an in-file marker: 2». Files hidden by ignore show only in the diagnostics block, as ignore: 3 files next to the diff count. And binary files the bot drops silently: they are in neither the summary nor the block, only in the bot's log.

The run has findings, but no new threads. The summary then shows «Findings: N» with «new inline comments: 0», and says why on a separate line or in a folded block. Findings below the min_severity threshold go into the «🔇 Below the min_severity threshold» block (by default the threshold is info, and nothing falls below it). Threads that are already open are not duplicated, and ones you resolved are not raised again — the line «From earlier runs — already under review, not duplicated: N; resolved by you, not reopened: M» says so.

The review asked a question instead of making a finding. With questions: true, a candidate the judge dropped as unproven may come back as a ❓ thread. It is a question for the author, not a finding, and it counts in neither the counters nor the gate. The summary gives it a line of its own: «❓ Questions for the author: N (outside the severity gate — see the threads)». Questions only work with a judge configured. «No findings» and a question waiting for you do not contradict each other.

Next