Vibhanshu Sharma
active · powerplay
PORTFOLIO.SYS›content›blog›prs-04-orchestrator-and-cost.mdx
Markdown · 10 min read · 2026-07-28

We Put a Cheap Model in Charge of the Expensive Ones

Part 4 of 5 — Ten passes and two context layers on every push costs real money, and a reviewer that costs too much gets throttled into uselessness. Here's the orchestrator that decides what runs, the four levers around it, and the invariant none of them are allowed to break.


// tl;dr
  • Ten passes plus two context layers on every push got expensive enough to risk being throttled off — and a reviewer that's off isn't a reviewer.
  • A cheap triage model proposes which passes to run; deterministic code overrules it for anything touching auth, money, or migrations.
  • Model tiering, prompt-prefix caching, bigger diff chunks, and budget guards stack multiplicatively, not additively.
  • The invariant: savings come from where compute is spent, never from how much reasoning is allowed.

Part 4 of 5 — How I built PRS at Powerplay. Part 3 gave the reviewer a map of the whole repository. Everything up to here made it better. All of it made it more expensive.

By this point a single review meant roughly ten parameter passes, a whole-repo symbol map, and a per-PR definition resolver — running on every push, across four repos.

That creates a failure mode that has nothing to do with code quality: a reviewer that costs too much gets throttled. Someone limits it to once per PR, or to certain branches, or turns it off for the team that ships most. And a reviewer that isn't running isn't a reviewer.

So cost work here wasn't about saving money. It was about keeping the thing switched on.


Measure before optimising

I refused to guess, and I also deliberately skipped a paid synthetic bake-off — spending real money to prove savings on fake traffic tells you nothing, because your PRs have a shape and that shape is the entire variable.

I read the number off live pull requests instead:

~$2.55 per full review, before any routing.

Then I made the measurement permanent, which mattered more than the number itself. PRS posts a per-pass cost ledger on every single review, so a regression is visible immediately rather than at the end of a billing month.

Here's a real one — click any row for why that pass sits on the tier it does:

Per-pass cost ledger — posted on every reviewclick a row
PassModelInputCachedOutputCost
▸Deep Reviewstrong-tier8,8510279$0.0263
▸Cross-Filestrong-tier5,921027$0.0152
▸Security · Tenancycheap-tier11,1360554$0.0145
▸Conventions · Archcheap-tier10,147087$0.0107
▸Testing · Hygienecheap-tier8,835091$0.0094
▸Quick Scancheap-tier8,576085$0.0091
Total53,46601,123$0.0851
Two levers are visible right in this table. The model column is the first — two deep passes on the strong tier, four breadth passes on the cheap one. And the cached column reads 0 because this was a first review; on a re-review most of that input arrives at roughly a tenth of the price.

You cannot route spend you can't see. The telemetry came first; every lever below was chosen by reading that table.


The orchestrator

The biggest lever isn't which model runs a pass. It's whether that pass runs at all.

Running every pass on every diff is wasteful — a data-layer specialist reviewing a CSS change produces confident opinions about nothing, and you pay full price for them. But running too few is dangerous in a way that doesn't show up until something reaches production.

So a cheap triage model reads the diff first and proposes which passes to run and at what depth. Then deterministic code decides whether it's allowed to.

Pick what landed in the PR:

Orchestrator — pick what landed in the PR2 / 8 passes
quick_scan
ALWAYS RUNS
security_tenancy
skipped
data_layer
skipped
performance_async
skipped
testing_hygiene
skipped
conventions_arch
triage chose
deep_review
skipped
cross_file
skipped
Nothing here can break auth, corrupt data or leak money. Triage skips the passes that would only generate noise, and code has no reason to object.
Let a cheap model propose. Let deterministic code constrain. Triage can make a review cheaper. It can never make a money-path review shallower.

The two scenarios that define the design

The auth and triage errored scenarios are the two that define the design.

In the auth case, triage sees a small, unremarkable diff and proposes a cheap review. It's not wrong about the size — it's wrong about the stakes, which a diff's shape doesn't tell you. Code overrules it: anything touching auth, money, DB queries or migrations gets security, data and deep, regardless of what the model thought. Backend repos always get the full set. Quick Scan always runs, on everything.

In the error case, the orchestrator being unavailable falls back to the full review, not the cheap one. This is a deliberate asymmetry: an unavailable opinion must never be read as "nothing to worry about here." That's the same fail-closed instinct that shows up again, painfully, in Part 5.

Propose cheap, constrain in code

The pattern generalises, and I've reused it three times since:

Let a cheap model propose. Let deterministic code constrain.

The model's judgment is an input to the decision, never the decision. Triage can make a review cheaper. It can never make a money-path review shallower.


Four more levers, and why they multiply

1 · Model tiering. Quick Scan, the specialists, and all incremental-push passes run on a frontier model roughly 54% cheaper per token. The strongest model is reserved for Deep Review and first-activation depth, where extra reasoning genuinely changes which findings come out. It's an Action input, so tuning cost is config, not a code change.

2 · Prompt-prefix cache. The provider discounts the longest unchanged prefix of a request by about 90%, so the prompt is ordered most-stable → least-stable under a per-pass cache key.

This one is worth playing with, because the failure mode is invisible until you look at a bill. Try moving the diff up:

Prompt order — the cached prefix
1.
parameter prompt + base rules
identical per slug, every PR — the anchor
~90% OFF
2.
PR context
stable across a PR's pushes
~90% OFF
3.
learnings (semantic re-rank)CHANGES PER PUSH
can shift between pushes
full price
4.
prior findingsCHANGES PER PUSH
accumulate each run
full price
5.
## REFERENCED DEFINITIONS
per-PR (Part 3)
full price
6.
the diffCHANGES PER PUSH
largest, and changes every single push
full price
SHARE OF PROMPT AT THE DISCOUNTED RATE36%
This is the shipped order. Most-stable first, the diff always last. The provider discounts the longest unchanged prefix, so every block you place before a volatile one keeps its discount — and the diff, which is both the largest block and the one that changes every push, can only ever be last.

3 · Chunk sizing. I raised the diff chunk size to 250KB. Fewer calls on large PRs — and because the stable prefix makes chunks 2..n roughly 3× cheaper than chunk 1, halving the chunk count saves more than the bigger chunks cost.

4 · Budget guards. A per-PR cap on the expensive full-review path, while incremental push reviews stay unlimited and cheap. Rapid-push storms dedup via git patch-id, so force-pushing the same tree doesn't re-bill a full review.

They stack multiplicatively rather than additively, which is the whole point. A typical re-review is a cheaper model running fewer passes over a mostly-cached prompt in fewer chunks. Each lever is unremarkable alone; together they're the difference between "runs on every push" and "runs when someone remembers."


The invariant

Cutting cost usually means cutting quality. It doesn't here, and the reason is one rule I held everything to:

The savings come from where compute is spent — never from how much thinking is allowed.

Forced high reasoning effort, the verify gate, the rules-of-evidence contract, the code-enforced severity floors — all of it applies to every pass, on every tier, in every scenario above. Not one lever lowers the bar a finding has to clear. They only change which specialists show up and how much of their prompt was already paid for.

That distinction is what let me be aggressive about cost without ever having the conversation where someone asks whether the reviewer got dumber, and I don't have a confident answer.


What it bought

The point of all of this was never the bill. It was permission to run the expensive version everywhere, all the time, instead of rationing it to the PRs someone judged important — which is exactly the judgment call an automated reviewer exists to remove.

PRS is now honest, well-fed, and affordable enough to leave on.

What it isn't, yet, is finished. Accuracy isn't a milestone you ship; it's a thing that decays unless something keeps pushing on it — and that turned out to need a feedback loop, an adversarial gate, and a genuinely humbling week where the bot took apart the guardrail I'd built to keep it honest.

Next: Still getting more accurate →

Keep reading

We Shipped an AI Code Reviewer With Three Prompts. It Was Wrong Too Often and Quiet Too Long.

2026-07-28 · 9 min read

One Reviewer, Four Codebases, Four Different Definitions of Correct

2026-07-28 · 10 min read

Our Cross-File Pass Couldn't See Other Files. Tree-sitter Fixed That.

2026-07-28 · 10 min read

Accuracy Isn't a Milestone. Here's What Keeps Pushing On It.

2026-07-28 · 11 min read
← all posts