Cloud models are banned: review in a closed loop
Some code never gets to leave the perimeter. If your security policy bans cloud models, or there is simply no way out, the review can run entirely on your own hardware. The product works with the OpenAI-compatible API, so Ollama or vLLM on your own server is a configuration, not a fork.
How good such a review is depends on the model you run, not on where it runs. We ran four open-weight models, from 122 billion parameters to 2.8 trillion, over a diff with four known defects. On average they found between 1.7 and all four, and the largest found all four in every run. The cloud pair from the rest of this section finds three on the same diff.
The rest of this page: which model to run and on what hardware, how to set up Ollama and vLLM so that they neither cut the prompt nor call home, what to show your security officer, and why the perimeter is held by the installation rather than by the repository.
Which model to run
The measurement is from 24 September 2026. The diff is the coupon commit in our sample repository (the same commit on GitLab): 31 lines and four real defects — fractional cents, a customer's e-mail in a log line, tax computed before the discount, and a coupon lookup on a plain object that also reaches inherited prototype properties. One model plays both roles, generator and judge, which is the usual shape of a closed loop. Three runs per model.
| Model | Parameters, total / active | Hardware, roughly | Findings | Of 4 defects |
|---|---|---|---|---|
| Qwen3.5-122B-A10B | 122B / 10B | 1–2 × H200 | 2.7 | 2.3 |
| DeepSeek-V4-Flash | 284B / 13B | 2 × H200 | 2.0 | 1.7 |
| DeepSeek-V4-Pro-0813 | 1.6T / 49B | 8 × H200 | 3.3 | 2.7 |
| Kimi K3 | 2.8T / 104B | 16 × H200 | 5.3 | 4.0 |
| for comparison: claude-sonnet-5 + judge claude-opus-5 | — | the vendor's cloud | 3.7 | 3.0 |
We ran the open weights through OpenRouter: we do not have a GPU rack. The weights are the ones you would load into vLLM, but OpenRouter picks the provider, and some providers serve quantized weights — on full precision your numbers may differ. Time and price belong to the hosting provider and would be different for you, so they are not in the table. The hardware column is a rough fit by the size of the weights, with room for the context. The cloud pair row is the measurement of 7 September on the same diff, from which model finds, which one judges.
The largest open model outdid the cloud pair on this diff: Kimi K3 found all four defects in every run. DeepSeek-V4-Pro came close to the pair, but never found the prototype lookup. We found no false finding in what the four published: everything beyond the four known defects was real. For example, an unknown coupon silently gives a zero discount while the log says «applied»; a discount and a coupon together can push the total below zero.
The open models get the line numbers wrong: usually by one to three lines, DeepSeek-V4-Flash by twenty and more; the cloud pair's are exact. That is not cosmetic. The judge checks a finding against the line it names, and in two runs out of twelve it dropped a correct finding because it read the wrong code. In one of them the judge explained just that: line 51 is safe, the violation is on line 53 — the defect was on line 52.
Connecting your own model
The product does not tell a cloud from your own server: Ollama, vLLM and other servers with the OpenAI-compatible API look the same to it. Name the models explicitly, though. Without roles in your personal config the CLI takes them from the repository's .reviewgate/config.yml, and if a cloud model is named there, it asks your server for it — we got a 404 exactly that way.
llm:
provider: ollama
base_url: http://llm.internal:11434/v1
timeout_ms: 1800000 # 30 minutes; the default is 10
generators:
- model: qwen3.5:122b-a10b
judges:
- model: qwen3.5:122b-a10bIf the server is Ollama, two daemon settings are mandatory:
OLLAMA_CONTEXT_LENGTH— the context size. Without it Ollama picks the context itself, by the amount of video memory: on our machine it came to 4096 tokens, while the review prompt for a 31-line diff weighs about 7,600. Ollama cuts the excess silently — the model sees neither the rules nor the instructions, and the review still «succeeds». The only trace is the linetruncating input promptin the daemon's log.OLLAMA_NO_CLOUD=1turns the cloud features off. Without it the daemon keeps a connection to ollama.com, and the models tagged:cloudin its library run on Ollama's servers, not on yours.
$ sudo systemctl edit ollama
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=64000"
Environment="OLLAMA_NO_CLOUD=1"
$ sudo systemctl daemon-reload
$ sudo systemctl restart ollama
# after a review: no output means nothing was cut
$ journalctl -u ollama | grep "truncating input prompt" For the bot it is the same, as environment variables. The roles in a repository's .reviewgate/config.yml win over them: if the repository names a cloud model, the bot asks your server for it and the review fails with an error. For repositories without roles of their own, LLM_MODEL sets the generator and LLM_JUDGES the judge.
LLM_PROVIDER=vllm
LLM_BASE_URL=http://llm.internal:8000/v1
LLM_MODEL=Qwen/Qwen3.5-122B-A10B
LLM_JUDGES=Qwen/Qwen3.5-122B-A10B
LLM_TIMEOUT_MS=1800000 If the server is vLLM, truncation is not a worry: it does not cut an overflowing prompt, it answers with an error, and the review says so in the merge request. It has telemetry of its own, though: by default vLLM sends anonymous usage stats — the hardware, the model's architecture, the settings, no prompts. VLLM_NO_USAGE_STATS=1 in the vLLM server's environment turns it off.
The timeout is not a formality. By default a model call waits 10 minutes. In our measurement the Qwen3.5-122B judge reasoned for 25 minutes in one run and produced 131 thousand tokens, and the two attempts before that were cut off by the timeout, at 15 and at 30 minutes. Give the calls room, and see how long a run takes before you put the review on a hook.
Show your security officer, do not tell them
A closed loop is a claim until somebody checks it. Below is the full list of what the model receives during a review, and the commands that let a sceptic see for themselves where the connections go.
What the model receives besides the product's own instructions:
- the diff of the change, and the paths of the files excluded from the review (by
ignoreor by a marker inside the file), without their contents; - your team rules, the team's allowances from
dont_flag, and thereview_prompttext — all from.reviewgate/config.yml; - from the bot, the merge request's metadata: title, author, branches, project, link, and the description up to 2,000 characters; the CLI sends none of it;
- the judge gets the whole files the findings touch and the definitions they import; with
full_file_contextthe generator gets the same for the changed files; - with
environment_context, the stack versions read from your manifests; - in reply mode, the thread being answered and the lines of code around the one discussed.
What the model never gets: your git history, the working tree beyond those files, other repositories. And the product does not store the code anywhere after the run.
And the check. Run a review from the CLI against your model and, while it runs, look at the network connections of two processes: the review and the model.
$ lsof -nP -i -a -p "$(pgrep -d, -x reviewgate)"
COMMAND PID FD TYPE NAME
reviewgat 84624 14u IPv4 TCP 127.0.0.1:53958->127.0.0.1:11434 (ESTABLISHED)
$ lsof -nP -i -a -c ollama -c llama-server
COMMAND PID FD TYPE NAME
ollama 84616 3u IPv4 TCP 127.0.0.1:11434 (LISTEN)
ollama 84616 7u IPv4 TCP 127.0.0.1:11434->127.0.0.1:53958 (ESTABLISHED)
ollama 84616 10u IPv4 TCP 127.0.0.1:54060->127.0.0.1:53960 (ESTABLISHED)
llama-ser 84745 3u IPv4 TCP 127.0.0.1:53960 (LISTEN)
llama-ser 84745 4u IPv4 TCP 127.0.0.1:53960->127.0.0.1:54060 (ESTABLISHED)
# the same daemon without OLLAMA_NO_CLOUD=1 adds one line (34.36.133.15 is ollama.com):
ollama 84446 11u IPv4 TCP 10.8.1.1:53824->34.36.133.15:443 (ESTABLISHED) That is live output from a real review on a local model: the review talks only to the model's port, Ollama only to its own llama-server, and all of it stays on 127.0.0.1. Without OLLAMA_NO_CLOUD=1 the same command showed one more line — the connection to ollama.com. For a bot in a container, docker inspect shows its networks and environment — where it is set up to go — and tcpdump on the host shows where it actually goes. The strongest check is an egress allowlist that simply has no line for a model vendor: if the review still works with the allowlist on, the argument is over.
Building the loop
The installation holds the perimeter, not the repository. The bot reviews a merge request by the .reviewgate/config.yml from that merge request's own branch, and a role in that file may name any of the bot's backends. If the named backend is not there, the main generator goes to the installation's main provider. Which means that if the installation has a cloud anywhere — as a named backend or as the main provider — one line in a merge request sends its diff out, and the summary reports it only afterwards. Code that must not leave therefore needs its own installation with no cloud at all, behind an egress allowlist without model vendors. Splitting by repository on one installation is only for code that may leave anyway.
Make the strongest model you have the judge. The line of a finding drifts at the generator, and then the judge decides: a weak one reads the wrong code and drops a correct finding, a strong one finds the right line itself. In such cases the DeepSeek-V4-Pro judge kept the finding, and in one run it named the real line numbers outright. If the big model cannot carry every run end to end, let a smaller one generate and keep the big one as the judge.
And spend on the model, not on passes. Recall moves with the model itself: extra judge passes cannot add what the generator did not see, because the judge is only allowed to take away. In the table above the gap between the weakest and the strongest model is more than two defects out of four.
When it goes wrong
The findings are strange: about code that is not there, on lines where it is not. Most likely Ollama has cut the prompt, and the model saw only its tail, without the rules and the instructions. Look for the line truncating input prompt in the daemon's log and raise OLLAMA_CONTEXT_LENGTH.
The review fails with a 404 and the name of a cloud model. The roles came from the repository's .reviewgate/config.yml, and a cloud model is named there. For the CLI, set the roles in your personal config; for the bot, change the roles in the repository or remove them, so that LLM_MODEL and LLM_JUDGES take over.
The judge times out, and the findings are published unchecked. The run does not fail: the judge drops out, the findings go out without the second pass, and the summary says so in a line of its own. The model answers slower than the call waits — give timeout_ms (LLM_TIMEOUT_MS for the bot) more room.
Ollama shows a connection to the outside in lsof. The variable OLLAMA_NO_CLOUD=1 is not set, and the daemon keeps a connection to ollama.com. Add the variable to the service's environment and restart the daemon.
The review reports that the context does not fit the model's window. The prompt did not fit into the window vLLM was given with --max-model-len: vLLM does not cut such a prompt, it answers with an error, and the product passes it on with advice — generated files out through ignore, full_file_context off, or a smaller change. You can also raise the window itself if the server has the memory for it.
The judge drops correct findings, pointing at a line number. The reasons show in the 🔬 block with diagnostics: true and in the reviewgate review --json report. The lines drift at the generator, and the judge is too weak to find the right one itself. Make a stronger model the judge — see the chapter above.
Next
- Models and providers— every setting of the schema, and named backends
- Which model finds, which one judges— why the judge may only take away, and why it must be stronger than the generator
- Security— the same list from the vendor's side: what leaves and what is stored