The Judge Tax: Why Cheap-Model Routing Needs 23% Offload, Not 4%

Share
The Judge Tax: Why Cheap-Model Routing Needs 23% Offload, Not 4%

Routing to a cheaper model only pays if the judge costs less than the savings it unlocks. Here is the formula, the two designs that produce wildly different break-even points, and the term that quietly dominates both.

The first mistake teams make with AI routing is treating it like a trick. It is not a trick. It is an economic trade, and like any trade it has a spread. You pay a judge to decide where each request goes, and you hope the savings on the requests it sends downstream cover what the judge cost you.

Most routing write-ups skip the spread entirely. They show you the cheap model's price next to the expensive model's price, wave at the gap, and call it a strategy. That gap is not your saving. Your saving is what is left after the judge takes its cut.

Call that cut the Judge Tax. This post is the math for it, including the part that most write-ups get wrong.

The number that starts the argument

LangChain benchmarked exactly this on their Deep Agents suite: 145 multi-step agentic tasks, averaging 6.3 model calls each, across support dialogue, incident investigation, and workflow automation.

Configuration Accuracy Cost per run
Claude Opus 4.8 alone 86.0% $11.45
Routed (Opus + Nemotron 3.5 Lightning) 80.0% $3.00
Nemotron 3.5 Lightning alone 77.7% $0.72

Two findings matter more than the headline 74% cost reduction.

The first: only 7% of agent turns needed the frontier model, but those 7% consumed 68.4% of total spend. That is the asymmetry the whole field is chasing. A small slice of your traffic is eating most of your bill.

The second is the one nobody quotes. In the routed configuration, the judge consumed 21.2% of spend. That is about $0.64 of every $3.00 run. The judge was not free. It was the second-largest line item in the architecture that was supposed to be saving money.

That is the Judge Tax, measured in the wild.

The formula

Let:

  • E = cost of the expensive path, per request
  • C = cost of the cheap path, per request
  • J = cost of the judge, per request
  • p = fraction of requests the cheap path successfully absorbs (your offload rate)

If the judge classifies each request before anything runs, your routed cost is J + p·C + (1−p)·E, against a baseline of E. Routing wins when:

p > J / (E − C)

Minimum Offload = Judge Cost / (Expensive Cost − Cheap Cost)

Let's put real August 2026 list prices in it. Take a support-triage request at roughly 2,000 input tokens and 300 output tokens:

Path Model Price (in / out per Mtok) Cost per request
Expensive Claude Opus 4.8 $5 / $25 $0.01750
Cheap Claude Haiku 4.5 $1 / $5 $0.00350
Judge Gemini 3.1 Flash-Lite $0.25 / $1.50 $0.00061
p > 0.00061 / (0.01750 − 0.00350)
p > 0.00061 / 0.01400
p > 4.3%

Offload 4.3% of your traffic and you are ahead. That is a low bar, and it is why routing looks like free money in a spreadsheet.

Now watch it break.

The same models, a different design, a 5× harder bar

The formula above assumes the judge decides first. That is classifier routing: look at the request, pick a lane, run one model.

But most teams do not build that. They build escalation routing, because it is more accurate and more intuitive: let the cheap model try, have a judge grade the answer, escalate to the frontier model only if the grade is poor. Switchyard calls this cascade routing. It is the pattern in nearly every "local-first agents" post you have read, including two drafts of my own.

In that design you pay for the cheap attempt on every single request, including the ones that end up escalating. The cost becomes C + J + (1−p)·E, and the break-even moves to:

p > (C + J) / E

Same three models. Same prices. Different answer:

p > (0.00350 + 0.00061) / 0.01750
p > 23.5%

4.3% versus 23.5%. The identical model lineup needs more than five times the offload rate to break even, purely because of when the judge runs. Now add retries, the "two strikes before escalating" rule that shows up in most escalation designs. It gets worse:

Retry rate on first-pass failures Minimum offload to break even
0% 23.5%
15% 27.0%
30% 30.5%

If you have been quoting J / (E − C) at a system that actually runs the cheap model first, you have been underestimating your break-even by a factor of five. That formula is not wrong; it is just answering a different question than the one your architecture is asking.

Check which design you built before you trust either number.

The term that dominates both

Here is where the arithmetic stops being about tokens.

A routed system is less accurate than a frontier-only one. The LangChain benchmark lost six accuracy points. Those failures do not evaporate. A human picks them up.

Price that human. Substitute your own figures here; ours are placeholders to show the shape. Three minutes of cleanup at a $40/hour loaded rate is $2.00 per corrected request:

  • At a 5% correction rate: $0.10 per request, or 5.7× the entire Opus path
  • At a 2% correction rate: $0.04 per request, or 2.3× the entire Opus path

And the marginal version, which is the one to write on the wall:

One point of accuracy costs $0.020 per request in human cleanup. The entire per-request saving from routing is $0.014.

One accuracy point is 1.43× your whole routing margin. A design that saves 80% on inference and costs you two points of accuracy is not a cost reduction. It is a cost transfer, from a line item your CFO can see into one they cannot.

This is not an argument against routing. It is an argument against evaluating routing on the inference bill alone.

So when does it actually win?

Look back at the benchmark, because it cuts the other way and the honesty matters.

Routing saved $8.45 per run there. The six-point accuracy drop, priced at $2.00 of cleanup per failed run, costs $0.12. Routing wins by roughly seventy to one.

Why so decisively, when the per-request math above was marginal? Because those runs averaged 6.3 model calls at high token volume. The judge is a fixed cost paid once per decision; the savings scale with how much work sits behind that decision. Long agentic runs amortise the Judge Tax across many expensive calls. A single 300-token classification does not.

That gives you a clean rule:

Routing pays in proportion to how much expensive work each judge call is deciding the fate of.

Workload shape Verdict
Long agentic runs, many calls, large context Routing wins comfortably
High-volume single-shot classification Marginal. Check the design regime first
Low volume, any shape Not worth the operational cost
High stakes, human reviews everything anyway Skip routing, cut the review cost instead
Deterministic and rule-expressible Skip the LLM entirely

That last row deserves more attention than it gets. A regex, a lookup table, and a validation rule cost nothing per request, never drift, and never need a judge. A meaningful share of what teams route to a "cheap model" should not be hitting a model at all.

What Switchyard actually is

If you are going to test this, do not build the routing layer yourself.

Switchyard is an Apache-2.0 routing proxy from NVIDIA's NeMo group, written in Rust. Its useful trick is protocol translation: it speaks both OpenAI and Anthropic API formats, so you can point an existing application at it and swap backends underneath without touching application code. It ships several strategies: random-routing for A/B splits and baselines, llm-routing for classifier-style decisions, and cascade for the signal-driven escalation pattern described above.

Configuration is YAML:

endpoints:
  openrouter:
    api_key: ${OPENROUTER_API_KEY}
    base_url: https://openrouter.ai/api/v1

targets:
  strong:
    endpoint: openrouter
    model: openai/gpt-4o
    format: openai
  weak:
    endpoint: openrouter
    model: openai/gpt-4o-mini
    format: openai

profiles:
  smart:
    type: random-routing
    strong: strong
    weak: weak
    strong_probability: 0.3

Start there, with random-routing rather than a classifier. A fixed probability split is the cleanest way to measure your actual E, C, and accuracy delta on your own traffic before you introduce a judge and start paying the tax. You cannot compute a break-even from numbers you have not measured.

pip install "nemo-switchyard[cli,server]"
switchyard launch claude --model openai/gpt-4o-mini \
  --api-key "$OPENROUTER_API_KEY" \
  --base-url "https://openrouter.ai/api/v1"

Once it is running, the metrics worth a dashboard are route accuracy, escalation rate, fallback rate, p95 latency, and the percentage of requests a human corrects after the first answer. That last one is the expensive one. Instrument it first.

What you give up

Routing is not free beyond the token math, and the costs are the sort that show up in month three.

  • Latency. The judge is a network hop with a model behind it. In escalation designs, escalated requests pay for the cheap attempt, the grade, and the frontier call in sequence. Your p95 gets worse even as your average bill improves.
  • Two systems to keep honest. Every prompt change now needs validating against both tiers. Skip it and the tiers drift apart quietly.
  • An eval harness you now have to maintain. Thresholds tuned on a labelled holdout set in March are not valid in September. Without periodic re-validation, your routing policy is a guess wearing a number.
  • Silent degradation. When a provider updates the weak model, nothing breaks. Accuracy just slides, the correction rate creeps, and the cost lands in a budget nobody is watching. This is the failure mode that gets caught last.
  • A judge that can be wrong in both directions. Over-escalation burns the savings; under-escalation ships bad answers. Both are invisible unless you log the disagreements between what the cheap path said and what the judge decided.

Self-hosting the cheap tier adds a standing operational chore on top of all of this. It is the same trade we wrote up for Inngest, with the same conclusion: worth it when the volume justifies it, a distraction when it does not.

The worksheet

Twenty minutes with a spreadsheet answers this before you write any code.

  1. List your top five request types by volume. Not by how interesting they are. By count.
  2. Measure E per request type. Actual tokens in and out at current list prices, not estimates.
  3. Measure C on the same requests. Run them through the cheap model. You need real token counts, because cheap models are often more verbose and eat some of the gap.
  4. Price the judge. Its input is the original request plus the candidate answer. That is bigger than people assume.
  5. Identify your design regime. Judge first → p > J/(E−C). Cheap model first → p > (C+J)/E. Get this right; it is the five-fold difference.
  6. Measure the accuracy delta on a labelled holdout set, then price it at your real cost of a human correction. Subtract that from the saving. This step is the one everyone skips and the one that decides the answer.
  7. Compare the required offload rate to what you can plausibly achieve. If the math needs 30% and your traffic is 20% routine, stop. You found your answer for the price of an afternoon.

If the routing policy only works on paper, it is not ready. If the fallback path is slower than the business can live with, the threshold is wrong. And if the cheap path is not dramatically cheaper, meaning C sits anywhere near E, no judge is cheap enough to rescue the trade.

FAQ

What is the Judge Tax? The per-request cost of the model or logic that decides where each request should go. It is subtracted from your routing savings, and in LangChain's benchmark it accounted for 21.2% of total spend in the routed configuration.

What is a good offload rate? There is no universal number. That is the point of the formula. With Opus 4.8 against Haiku 4.5 and a Flash-Lite judge, classifier routing breaks even at 4.3% and escalation routing at 23.5%. Compute yours from your own token counts.

Should I fine-tune a model to improve routing? Almost certainly not yet. LoRA, SFT, and GRPO are standard playbooks now, but fine-tuning adds dataset curation, training runs, regression risk, and deployment overhead. It also reliably masks problems that were actually bad prompts, missing evals, or a badly drawn routing boundary. Fix those three first. Most teams never need step four.

Does routing require running models locally? No. Local inference is one way to make C small, and it is the most effective one when volume is high, but the math only cares about the price of the cheap path. A cheaper hosted model gets you the same trade without the operational burden.

Is Switchyard a Practical Works product? No. It is NVIDIA's open-source project, Apache 2.0. We do not maintain it and have no affiliation with the project.


Cheaper per call is not the same as cheaper. The gap between those two sentences is where migration budgets go missing, and it is wide enough to measure before you commit to it.

Do the arithmetic before the migration, not after.

If you want a second pair of eyes on it, bring your measured E, C, and J, your design regime, and your current correction rate. We will work through whether the cheap path actually wins before you spend another quarter on the wrong model tier.

Hire us to build it →

Prices cited are published list prices as of August 2026: Anthropic and Google. Benchmark figures from LangChain's Switchyard agent routing benchmark.

Read more