Your review got expensive: where the tokens actually go
Every call to the model carries the engine's own text ahead of your code: the preset, the team rules, the guidance, the contract. On our sample repository that is 12 365 characters, and they are the same for a one-line change and for a merge request touching a hundred files. The diff is the variable part: on the coupon commit (the same commit on GitLab) it is 2 847 characters. That is the first thing to know about the bill: the size of the constant part is set by your config, not by your code.
The second is that the same review, run twice within an hour, costs different money: the first run pays for writing the cache, the repeat for reading it.
The third is that two settings change the bill several times over: one multiplies it without mentioning it anywhere, the other divides it — and it is not the one people usually expect. This page is about all three, with the numbers from our own runs.
Almost every number on this page comes from our public sample repository (the same on GitLab): you can clone it and run the same reviews yourself. On your project the numbers will differ, but the bill is built the same way — from the same parts, with the same cache and the same settings. And you look at it in the same place: the run's diagnostics block, covered below. Prices are per the vendors' price lists on 23 September 2026.
What goes into the model, and how much it weighs
Measured on the sample repository — a nestjs preset, six team rules and the team's own review_prompt, no dont_flag assumptions, full-file context and environment context on. In total 12 365 characters:
| Part of the input | Size, characters | How it is paid for |
|---|---|---|
| preset instructions | 3 527 | every call, cached |
| default guidance | 3 056 | every call, cached |
| review contract | 2 969 | every call, cached |
team rules + dont_flag | 1 500 | every call, cached |
| service lines: block headings, separators, the language line, the rule-id instruction | 442 | every call, cached |
the team's own review_prompt | 429 | every call, cached |
| the project environment: versions from the manifests | 231 | every call, cached |
| the output-language directive | 211 | every call, cached |
| the diff itself | 4 086 | every call, never cached |
Everything above the diff stays the same between runs, so it is cached. The diff is not cached and cannot be: it is different every time. In tokens, which is what the bill counts, the run's diagnostics block on the coupon commit shows the same picture: of 14 432 input tokens, 8 517 are the constant part (the call line shows it as «cache: 8.5K»), and the other 5 915 are the diff, whole files and service lines. And the constant part grows with the config: for a team with two dozen rules like the ones in our sample, the rules block alone outweighs the preset.
And it is paid per role, not per run. That whole system block goes to the generator, then to every judge, then to the arbiter — each of them re-reads your rules and your preset. A schema with a generator, two judges and an arbiter pays for that text three times, and four times when the judges disagree and call the arbiter. This is the arithmetic behind «a sprawling rulebook is paid for on every call of every run» from the rules recipe. The judge's own input is heavier still: it gets the generator's entire system prompt, plus its own contract, plus whole files.
Why the same review costs different money
The constant part of the input is the same in every run, and that is exactly the part that can be cached: written once, then read on the following calls for a fraction of the price. At Anthropic, writing costs twice the input rate and reading a tenth — twenty times less for the same text. The diff is never cached: it is different every time. On the coupon commit the same run cost $0.12 with the cache write and $0.09 with the read, and the whole difference was that block. Vendors handle caching differently: Anthropic caches what the request marks and bills writes and reads separately; OpenAI-compatible providers cache on their own, by a matching start of the request; intermediaries such as OpenRouter have rules of their own, and for the same model the cache may not switch on at all. The check is simple: in the second run in a row, the call in the diagnostics block should show «cache: …» about the size of the system block. No number, no savings.
Any edit to any piece of the system prompt from the table above gives a new cache key — a one-word fix in one rule included. Edit a rule in the middle of the day, and the first run after that pays for the write again. The TTL itself lives in no config, neither the repository's nor the home one: it is the environment variable LLM_CACHE_TTL — in the bot's environment for the bot, in your shell for the terminal reviewgate; one hour by default, with 5m and off as the alternatives. If runs come less often than once an hour, the one-hour cache only hurts: every run pays double for the write. Then take either 5m — a write at 1.25×, which pays off when the next run comes within five minutes — or off, at the normal rate with no surcharge.
Two settings that change the bill several times over
full_file_context is the most expensive switch you have. Off by default, and for good reason. Turned on, the generator receives whole files of everything in scope — not just the files where findings landed — plus their templates, styles and imported definitions. That has its own budget of 200 000 characters, which does not replace the diff budget of 400 000: they add up.
Our own first hook run on the sample repository is what this looks like when it goes wrong: a diff of 37 lines across four files cost $0.33, because pnpm-lock.yaml went into the context whole — 143 962 bytes out of 146 179, that is 98.5% of the file context the model received. With the lockfile in ignore the same commit came to $0.01; the full story is in the hook ladder. The order therefore matters: put lockfiles, snapshots, generated code and fixtures into ignore, and do it before you turn full context on. Otherwise you pay for reading files that contain not a single line written by a person.
Only two settings remove a file before the model is called: the ignore list and the reviewgate-ignore-file marker with a reason inside the file itself. Everything else — dont_flag, min_severity, the rules — is text in the prompt or processing of the answer, and none of it makes the run cheaper. The engine itself also drops deleted files, empty diffs and binaries, but those are not your levers.
Turn on two settings that are off by default
# .reviewgate/config.yml
diagnostics: true # the 🔬 block: every call, its tokens, its duration
cost:
show: true # the price line under the summary
currency: "$"
models: # a price for EVERY model of the run, or no amount is shown
# $ per 1M tokens, vendors' price lists as of 2026-09-23
claude-sonnet-5: { input: 2, output: 10 }
claude-opus-5: { input: 5, output: 25 }
deepseek-v4-flash: { input: 0.30, output: 1.20 } # peak rate; off-peak is half
ignore: # one of the two levers that remove a file BEFORE
- "pnpm-lock.yaml" # the model is called; the other is the in-file marker
- "**/*.generated.ts"
- "**/__snapshots__/**"diagnostics adds a block to the report showing every model call: which role, which model, tokens in and out and how much of the input came from the cache, how long it took, how many findings the generator produced, how many the judge dropped or downgraded, and why. Without it a run is a black box with a number at the end.
cost.show adds the price line — and it needs one more thing: prices for every model of the run. Miss one and you will not see the money: in the terminal the line does not appear, and in the merge request summary it stays without an amount, naming the model that has no price. This is not hypothetical: three of our benchmark runs printed no price in the terminal because DeepSeek was missing from the price list, and the report was right to stay silent rather than show a number that counts only part of the run.
Two cases when a run suddenly costs more
The answer hit the output ceiling. The model was cut off mid-answer, the run repeated the call at a lower effort, and both calls are on your bill — one of them for nothing. The report says so plainly. On our own repository, on a branch of three recipe pages, this cost $0.94 instead of $0.52: back then the default ceiling was 16 000 tokens, now it is 64 000, but you can still set it lower yourself. OpenAI-compatible providers get no retry: a cut-off answer is a run error straight away.
A reasoning model turned out verbose. The same input on DeepSeek v4-flash gave 22.7K, 6.8K and 19.4K tokens of output across three runs — a threefold spread, for reasons on the model's side, not yours. Two of those three would not have fit into the former ceiling of 16 000 tokens. If your schema includes a reasoning model, its ceiling needs headroom that an ordinary one does not.
If you want a hard limit, there is a spend ceiling in tokens per run, off by default: LLM_BUDGET_TOKENS for the bot, and budget_tokens in the llm block of the terminal reviewgate's home config. It does not cut you off silently; it writes lines above the list of findings, like the other notices. At 80% it is a plain line: «backend anthropic is approaching the run budget: 41,000 of 50,000 tokens». At 100% it comes with ⚠️: «run budget for backend anthropic exceeded: 52,000 tokens against a limit of 50,000 — check the role layout and the size of the diff». At 300% the run stops at the nearest checkpoint: «🛑 review stopped by budget: … Not done: judging. What had been collected is published».
Next
- Which model finds, which one judges— what each role costs and adds
- Rules your linter cannot check— the block that is paid for on every call
- Config reference— cost, diagnostics, ignore