Your agent said «done» and never ran a review
The failure here does not look like a failure. You ask the agent to check the changes, it answers with a tidy list of remarks, and everything seems fine — except no review ran. The agent read your code itself and told you what it thought.
We repeated this on the sample repository: the same request — «check the branch before a push» — with three setups of the agent in Claude Code, one run each. Without the hint and without permission to call the review, the agent found our tools by itself and tried the review, was refused by the client, and reviewed the diff itself: four of the five known defects, one more than our review of the same diff. It mentioned that the review tool needed a permission it did not have and that its answer came from reading the code. With the hint but without the permission, it did the same, trying the review 20 seconds in and again at 54. With the hint and the permission, it got the verdict «gate failed» with three violations of the team's rules and added findings of its own on top. The agent's own reading can be strong, but it is not the team's review: without the permission no judge checked it and no gate ran, so the rule ids the agent quoted and its «so all four block» were its own reading of the rules, not a verdict. The records of every run — the requests, the tool calls, the agents' answers and the review report — are in the sample repository (the Russian measurement is on GitLab).
Three barriers stand between «the server is connected» and «the agent calls it»: the client has not picked the server up yet, the agent does not know when to call the tools or finds them late, and the agent is not allowed to call them. In our runs neither a refused call nor a late one failed the run: the session ended successfully, the refusal showed up as an error only inside the session log, and at best the agent mentioned it in its answer.
Before you start: you need reviewgate in PATH and a working review — your first local review covers that part. Everything below is about the agent, not about the tool.
Barrier 1. The client has not picked the server up yet
run reviewgate help agent-setup and do what it says The server is started by the client itself, with reviewgate mcp, from the client's working directory. Neither tool has a «repository» parameter: the review covers the git repository the client started the server in. Where and in what format that command is registered differs from client to client — for Claude Code it is .mcp.json at the project root, other clients have their own files. We deliberately do not guess other clients' formats: paste the line above into your agent's chat, and it will read the instructions and write the setup in its own client's format. Ready-made setups for Claude Code and Codex are in the agents reference.
The main trap here is that most clients read the MCP setup only at start-up. The agent wrote it in a running session, you asked it to check the changes — and it answers with its own reading of the code: it simply does not have the tools yet. Restart the session first. Some clients also ask whether to allow a new server from the project settings, and remember the answer. Decline, and the server does not connect, with the same symptom; your client's documentation says where to reset that answer.
Barrier 2. The agent does not know when to call the tools
Some clients, Claude Code among them, load tool descriptions lazily: at the start of a session the agent sees the tool names at best, not what they are for or when to call them. Ask it to check the changes, and it may read the diff itself, or get to the tools only after reading the code by hand. This is not a bug, and it is not fixed on our side: two paragraphs in your project's instruction file tell the agent which tools exist and when you expect them to be called. If you connected the server with the setup line, its instructions already had the agent add lines to the same effect; if you did it by hand, add these:
The standards of this project come from the mcp__reviewgate__get_team_rules tool —
call it BEFORE writing code, instead of reading the ADRs by hand.
Before finishing a task and before git push, check the changes with the
mcp__reviewgate__review_changes tool — it is the same judge and the same rules
that check the merge or pull request. Here is what our measurement showed. Asked to find out the team's standards, with the hint the agent answered from get_team_rules alone: 23 seconds, about $0.17. Without the hint it called the same tool, but also read the code by hand: 82 seconds, about $0.48. Asked to check the branch before a push, the agent reached for the review in every run: without the hint 51 seconds in; with the hint 20 seconds in, and — in the run that had the permission — only after nearly four minutes, because it read the code in the first 18 seconds and then reasoned about it on its own before calling the review. In the two runs without the permission, what stopped it was the next barrier. That is one run per scenario, so read it as what happened, not as a rate. The tool names in the block are in Claude Code's notation: in another client, bring them to its notation, or leave that to the agent by sending it the setup line.
Barrier 3. The agent is not allowed to call it
In both our runs without the permission, the agent tried the review, got the client's refusal and, eager to help, reviewed the diff itself. Both times it mentioned near the top of its answer that the review tool had not run; one of the two added that the gate had not been checked. You cannot rely on that: an agent can declare findings blocking from its own reading of the rules — ours wrote «so all four block». Within an agent session, the reliable sign of a real review is a review call — the review_changes tool, or reviewgate review run by the agent — with its answer in the session log; and the verdict is what the tool returned, not the agent's retelling: with the call in the log, our agent still put a finding of its own next to the judge's, under the team's rule and severity. It did say «the judge missed this one», but in the list of blocking findings the two look the same.
In an interactive session the client asks whether to allow the call, and allowing it once is enough. In a non-interactive run there is nobody to ask: the permission is granted up front, with a client setting or flag — in Claude Code that is --allowedTools, with an example in the agents reference. When it sets up the connection, the agent deliberately does not grant itself this permission: the decision is yours. The permission dialog is a guard against third-party code, and an agent that granted itself the permission would bypass it.
team:… — the tools work. If it retells your config in its own words, or says it will read the file, you are on one of the three barriers above.Two tools, and what comes back
get_team_rules returns the preset, the rules with their ids and severities, and, if one is set, your review_prompt — as data, not as prose. It is cheap, calls no model, and is meant to be called before writing code rather than after. The rule ids come back prefixed — team:money-in-minor-units — and that prefix is what lets the agent match a later finding to the agreement it breaks.
review_changes runs the review and returns the full report as JSON: the findings with file and line, the verdict, what the judge dropped, the token usage and the cost. The cost comes only if the team's price list covers every model of the run; otherwise it is null.
All the arguments are optional:
scope— what to review:uncommitted(the default), the uncommitted changes;staged, the index only; or an object with the base to compare against, such as{"base": "main"}, which covers what the merge or pull request will contain. A string such asmain..HEADis rejected.mode—full(the default), with the judge, orfast, the generator only: faster and cheaper.team_llm—trueto run with the roles from the team config rather than your home one.
Two things to know about the verdict. The exit code does not reach the agent: in the terminal a blocked run exits with code 2, while through MCP there is no exit code at all — the verdict lives in the verdict.gate field: pass, fail, or off when the team's gate is switched off.
And «clean» can be empty. If the agent has already committed its changes and calls the review without scope, there is nothing to review: the model is not called, and the report comes back with no findings and run.filesReviewed: 0. Branch on verdict.gate together with run.filesReviewed, and after a commit pass a scope with the base to compare against.
reviewgate mcp --fast The flags the client starts the server with set the defaults for calls. Scope, mode and roles — --staged or --refs, --fast, --team-llm — a call can override with its own arguments, while the other flags, such as --fail-on and --config, apply to every call. A team that wants cheap agent-side reviews puts --fast into the start command once, instead of hoping the agent remembers to pass mode.
The hook: a safety net if the agent did not call the review
Everything above is a request. A hook is not: it is a command started by an event rather than by the agent, so it does not matter whether the agent remembered the review. You have two such safety nets: a git hook on push, which works with any agent, and a hook of the agent client itself, if the client supports hooks.
A git hook on push is run by git, not by the agent, so it fires with any client — and on your own push too. It is the pre-push from the hook ladder: a full review with the judge before the code leaves the machine. It has one weak spot: an agent that was stopped can repeat the push with --no-verify and skip the check, and our own hook prints that hint, meant for people, when it stops a push.
The agent client's own hooks fire on its events — before a command runs, when a task ends. Not every client has them, and every client has its own format. Our reviewgate review --hook-stdin understands the Claude Code format, and for your own wiring it answers with an exit code under --hook-format exit. In Claude Code it acts on a git push in a Bash call — --no-verify included — and on the end of a task if you hang it on that event as well; everything else passes silently.
{
"hooks": {
"PreToolUse": [
{ "matcher": "Bash",
"hooks": [{ "type": "command",
"command": "reviewgate review --hook-stdin",
"timeout": 900 }] }
]
}
}Three properties of the hook worth knowing:
- It fails open. If the review itself breaks — no key, no network — the agent is not locked out of its own work, and the reason goes to stderr.
- It does not loop. When the client carries on because of the hook's own block, the repeated end of the task arrives marked, and the review does not run a second time. Otherwise «the agent fixes the finding → the task ends → the hook reviews again» would become an endless paid circle.
- The client can drop a long hook. Claude Code waits 10 minutes by default, and a dropped hook does not stop the push. That is why the example sets
"timeout": 900; on large diffs with an ensemble you can add--fastto the hook command.
One blind spot: the push detector reads the command word by word and does not recognise a push when git comes right after a quote in a shell string, as in bash -lc "git push". Everything else it sees: git push, /usr/bin/git push, git push --no-verify and even bash -lc "cd app && git push origin main".
How it is done in this repository
For a long time we worked without MCP. Our review of the agent's work ran on git hooks: a cheap one on commit, a full one with the judge on push.
The reason is that an agent does not always follow a rule in its prompt. The rule «docs are updated together with behaviour» lived in our instruction file for about a month and was still forgotten. Now three independent layers watch over it: the rule in the file, a skill with the procedure, and a hook that stops the end of a session once if behaviour was touched and the docs were not. Hence the pattern this whole page is built around: an instruction file helps the agent find the tool and understand when to call it, but only what fires by itself makes it follow the rule — a hook, not a request.
When it goes wrong
The agent brings its own reading instead of a review. That is one of the three barriers. Ask it the question from the «how to be sure» callout: the answer shows whether the tools are available at all.
The tools did not appear after the setup. Most clients read their MCP configuration only at start-up: restart the session. If that does not help, check whether you declined the server when the client asked for permission.
You passed arguments to get_team_rules and nothing changed. Its schema is empty and arguments are ignored silently — there is no way to narrow the rules it returns.
The hook did not fire on push. Two causes are possible. Either git came right after a quote in a shell string, as in bash -lc "git push", which is the detector's blind spot. Or the client dropped the hook on a timeout, and a dropped hook does not stop the push.
The same push was reviewed twice. Both hooks fired: the agent's hook before the command and the git hook at the push itself. If you do not need protection against --no-verify, keep only the git hook.
The review through the agent runs on other models than the team's. If your home config declares its own model roles, a local run — through MCP too — takes them whole instead of the roles in the team config, and the report warns about it. For the team roles, pass team_llm: true in the call or start the server with --team-llm. reviewgate doctor shows whose schema is in effect.
Next
- Your first local review— the tool itself, before the agent
- The hook ladder— a git hook on push, a safety net with any agent
- CLI, hook and MCP— ready-made setups for specific clients