LLM Gateway and Complexity Routing: Cutting AI Costs with Jev

Kuaray Deep Dive 01: how an LLM gateway plus a $0.00003-per-request classifier routes each prompt to the right model and cuts the bill by up to 62%.
Stop paying frontier prices for intern-level questions
Kuaray Deep Dive #01 — LLM gateway + Jev: how a ~$0.00003-per-request classifier decides which model each request deserves, and how much that saves.
The problem: one model for everything
Almost every AI product starts the same way: one API key, one model (usually the best one you can afford), and every call going to it. That works for a prototype. In production it becomes a bill that grows linearly with usage, and a good part of it is waste.
Real traffic is mixed. Someone asks "what are your opening hours?", someone wants an email summarized, someone needs a tightly coupled module refactored, and once in a while something actually requires heavy reasoning. Sending all of that to the top model is like hiring a senior engineer to answer the phone.
Vercel's guide on cost-aware routing puts it plainly: sending every request to a frontier model "is the simplest setup you can ship and the most expensive one you can run" [1].
The fix has two parts: an LLM gateway, and a classifier cheap enough not to eat the savings. That second part is where Jev comes in.
Part 1 — Why an LLM gateway, before any optimization
An LLM gateway is a single entry point between your application and the model providers. Your app calls one endpoint, and the gateway decides where the request goes and what happens to it on the way [2]. It pays for itself even before you add any smart routing:
- One API for every model. You switch providers or versions by changing a config string instead of code. Vercel AI Gateway, for example, takes identifiers in
creator/modelformat and handles the rest [1]. - Centralized credentials. Provider keys live in the gateway's vault instead of being scattered across services, so revoking access stops being a scavenger hunt [2].
- Reliability built in. Retries, provider failover, circuit breakers and timeouts live in one place instead of being reimplemented in every service [2]. A provider-specific failure "does not take the request down" [1].
- Real cost control. Per-team quotas, rate limits, spend caps and per-project cost attribution are all enforced at a single point [2]. On Vercel, each tier can get its own key with its own ceiling, and the gateway returns HTTP 402 once a limit is crossed [1].
- Caching. Prompt caching passes through to the provider, and identical requests are deduplicated [1][2]. This matters a lot, as the numbers below show.
- Free observability. Tokens, latency, model and status for every call, without instrumenting each service. You can see cost and latency per model and which model served each request [1][2].
- Governance and privacy. Zero Data Retention, "no training" and a pinned inference region (EU or US) are configurable at the gateway [3][14]. More on this in Part 7.
That's the foundation. Once it's in place, choosing the model per request stops being a refactor and becomes a routing rule.

Part 2 — The router is the hidden bottleneck
Complexity routing is a simple idea: classify each request into a tier (simple, medium, complex, reasoning) and send it to the cheapest model that can handle that tier.
The catch is who does the classifying. The usual options are:
- Heuristics (keywords, prompt length, presence of code). Cheap and instant, but brittle. Vercel's own guide starts here, with a list like
"debug","prove","refactor"[1]. - A small LLM as the classifier. More accurate, but now you're paying an LLM to decide which LLM to pay, and adding hundreds of milliseconds before the first token.
LLMs weren't built to classify. They generate text, and you parse the decision out of that text. That's the gap Jev fills.
Part 3 — What Jev is
Jev is a decision model from TypeSafe AI, released in early access on September 15, 2026 [4]. It doesn't generate text. You send a state (text up to ~32k tokens) and a set of typed questions, and it returns typed answers with probabilities [4][5]:
- Choice: picks one option from defined criteria (up to 255) and returns the probability distribution plus a confidence score [4][6].
- Score: places the state on an ordered scale of up to 10 levels [4].
- Noul: a calibrated yes/no that returns a probability between 0 and 1 [4][6].
TypeSafe calls it a "System One" model, a nod to Kahneman's fast thinking [4]. It makes no tool calls, holds no conversation and doesn't explain its reasoning. It answers the questions about the state in a single round trip [4].
Two things matter to whoever pays the bill:
- Price: $0.042 per million input tokens, and output tokens are free [7][8]. A typical call lands around $0.00001 to $0.00003 [6][8].
- Speculative fan-out: several questions about the same state go in one call. In MarkTechPost's guide, 10 questions batched together ran 7.1× faster and used 9× fewer tokens than 10 separate calls [6].
It's already available on the gateways that matter: Vercel AI Gateway (Sept 16, 2026) [3], LiteLLM, both as a pass-through [9] and as a native Auto Router classifier [10], and listed on OpenRouter and Cloudflare [4][11].
Part 4 — Jev as a routing classifier: the numbers
The LiteLLM team plugged Jev into their complexity router and compared it with Claude Haiku 4.5 doing the same job, across 80 cases, 240 calls and 4 tiers [10]:
| Metric | Jev | Haiku 4.5 |
|---|---|---|
| p50 latency | 127 ms | 688 ms |
| Expected-tier match | 95.0% | 73.8% |
| Cost per classification | $0.0000321 | $0.000827 |
| Difference | 5.4× faster | 96% cheaper |

An independent test by Classmethod, which reproduced NVIDIA NeMo Switchyard's 4-tier scheme, got 40/40 correct classifications at $0.000025–0.000027 per call, about 3× faster than the original Gemini Flash classifier [12].
The LiteLLM config is declarative:
- model_name: jev-router
litellm_params:
model: auto_router/complexity_router
complexity_router_config:
tiers:
SIMPLE: {{cheap_model}}
MEDIUM: {{mid_model}}
COMPLEX: {{strong_model}}
REASONING: {{frontier_model}}
classifier_type: jev
jev_classifier_config:
model: jev-latest
timeout_ms: 3000
circuit_breaker_enabled: true
Source: LiteLLM [10]. If Jev fails or times out, a circuit breaker (30 s default cooldown) sends everything to the default_model. The router never becomes a single point of failure.
Part 5 — Quantifying: how much it saves
Here's a concrete, reproducible scenario. It uses Anthropic's official pricing [13] and the classification cost measured by LiteLLM [10].
Assumptions
- 1 million requests per month
- 2,000 input tokens per request (1,500 of them a fixed, cacheable system prompt/context) + 500 output tokens
- Traffic mix: 50% simple · 30% medium · 15% complex · 5% reasoning
Tier mapping (price per million tokens, input/output) [13]
| Tier | Model | Input | Output | Cost/request |
|---|---|---|---|---|
| SIMPLE | Claude Haiku 4.5 | $1 | $5 | $0.0045 |
| MEDIUM | Claude Sonnet 5.5 | $2 | $10 | $0.0090 |
| COMPLEX | Claude Opus 5.5 | $4 | $20 | $0.0180 |
| REASONING | Claude Fable 5.1 | $10 | $50 | $0.0450 |
Monthly result
| Scenario | Cost/month | vs. all on Opus |
|---|---|---|
| Everything on Fable 5.1 | $45,000 | +150% |
| Everything on Opus 5.5 (baseline) | $18,000 | — |
| Gateway + Jev (routing) | $9,932 | −44.8% |
| Gateway + Jev + prompt caching | $6,861 | −61.9% |

How we got there:
- Routing: 0.5 × 0.0045 + 0.3 × 0.009 + 0.15 × 0.018 + 0.05 × 0.045 = $0.0099 per request, or $9,900/month. Jev adds $32.10/month (1M × $0.0000321), which is 0.3% of the bill.
- Caching: with 1,500 cached tokens read at 10% of the input price (5% on Opus 5.5, 2.5% on Fable 5.1) [13], the cost drops to $6,829 + $32 for Jev.
- If Haiku were the classifier: classification would cost $827/month instead of $32, and every request would wait ~560 ms longer before starting [10]. Total savings would fall from 44.8% to 40.4%, with worse tier accuracy.
Against the "everything on frontier" scenario (Fable 5.1), savings reach 84.8%.
Being honest about the math. The 50/30/15/5 mix is an assumption: measure your own traffic first (the gateway shows you). We ignored cache-write cost, which is charged once per cache window. Most importantly, routing only saves money if quality holds up in the cheaper tiers. That's why the next section exists.
Part 6 — The pattern that closes the loop: cheap first, escalate on evidence
Routing down carries an obvious risk: the cheap model gets it wrong. The answer isn't to give up on routing. It's to verify, and escalate only when needed. Vercel's guide already suggests escalating when the cheap answer fails a quality check [1], and Vercel lists "output verification with guardrails" among Jev's use cases [3].
The design we recommend:
- Jev
Choicepicks the tier on the way in. - The tier's model answers.
- Jev
Noulasks: "does this answer fully address the request?" It costs cents per million requests. - Below a probability threshold, the gateway resends to the next tier up.
- Gateway logs show the escalation rate per tier, so you tune the boundaries with data instead of intuition.
Because Jev's probabilities are calibrated in aggregate (0.8 means "right about 80% of the time for this kind of answer") [4], the threshold becomes an explicit business decision: how much quality risk you accept for how much savings.
Part 7 — Data custody: the prompt is personal data too
Everything that passes through the router is user content: names, emails, tickets, contracts. If you serve customers in the European Union (or handle data under Brazil's LGPD), where inference runs stops being a technical detail. The gateway is the right place to enforce that rule, because it already sees every request.
What the direct API doesn't solve. Today Anthropic's API offers only two inference geographies: global (default) and us, the latter at 1.1× pricing [15]. There's no way to keep inference in the EU by calling Anthropic directly. For Claude with EU residency, the path is a provider in a European region, such as AWS Bedrock (Ireland, Stockholm, Frankfurt via the EU cross-region profile) [16], behind the gateway.
What a gateway with regional inference does. Since July 2026, Vercel AI Gateway lets you pin the region per request [14][17]:
const { text } = await generateText({
model: 'anthropic/claude-sonnet-5.5',
prompt,
providerOptions: {
gateway: { inferenceRegion: { scope: 'zone', geoRegion: 'eu' } },
},
});
Three properties matter for compliance [14]:
- Fail-closed. If no provider can serve the request in the EU, it fails with HTTP 400. There's no silent fallback to another region.
- Proof, not promises. Every response reports where inference actually ran (
inferenceEndpoint.geoRegion). Log it and you have a per-request audit trail. - Per tenant. The region is chosen per request, so you can pin
euonly for European customers and leave everyone else on the cheaperglobalrouting.
What changes in the bill. On the gateway, Haiku 4.5, Sonnet 5.5 and Opus 5.5 are available in the EU at +10% per token. Fable 5.1 is US-only [18]. In our scenario with everything pinned to the EU, the REASONING tier falls back to Opus 5.5:
| Scenario (EU region) | Cost/month |
|---|---|
| Everything on Opus 5.5 (EU) | $19,800 |
| Gateway + Jev + caching (EU, REASONING → Opus) | $6,550 (−67%) |
The percentage savings look similar. But part of the "cheaper" comes from losing the strongest reasoning tier: that's a capability trade-off, not free savings.
The blind spot: the classifier reads the prompt too. Jev has to see the text to pick a tier. At the time of writing, TypeSafe doesn't publish where it processes data. Cloudflare's catalog lists zero data retention but no region [11]. If the rule is "nothing leaves the EU", you have four options:
- Check Jev's own region in the gateway catalog (the
regionsfield in/v1/models) and pineuon it too, if it's offered [14]. - Minimize what the router sees. The tier rarely depends on the full content. Send an excerpt with PII masked, or just structural signals. In LiteLLM, the classifier's context is configurable (default: 3 turns, 8,000 characters) [10].
- An in-EU classifier for regulated tenants: heuristics, or Haiku 4.5 pinned to the EU. It costs ~$900 per million requests instead of $32, still a small fraction of the savings.
- A self-hosted gateway in the EU. Vercel notes that it pins the provider's region, but your request may still pass through any Vercel region before being forwarded (region-pinned hosts are "coming") [14]. If that matters, run LiteLLM in a European VPC with Bedrock/Vertex in EU regions (in LiteLLM, use tag-based routing; the old
allowed_model_regionis deprecated) [19], or use a European gateway such as EUrouter (Amsterdam) [20].
The honest limits. Provider abuse monitoring may retain flagged requests outside your chosen region [14]. Zero Data Retention doesn't apply to every model (Fable, for example, isn't eligible) [14]. And, as Vercel itself points out, location is one GDPR obligation among many, not compliance on its own [14]. This isn't legal advice: check with your DPO. For the wider regulatory picture, see our guide to EU AI Act compliance for SaaS.

The caveats
- Early access. Jev is waitlisted and has no direct free tier, but Vercel AI Gateway offers a monthly credit on its free plan [7][8].
- Young benchmarks. The authors of the LiteLLM benchmark note that the labels weren't independently reviewed and that results depend on the prompts and tier definitions [10]. On TypeSafe's own dashboard, aggregate accuracy on general classification tasks is 67.8% versus 74.1% for the best comparator [7]. Run your own evals on your own traffic.
- Text only. No image, audio or video input [4].
- Not a generator. Jev decides, the LLM writes. The hybrid architecture is the point, not a limitation.
Conclusion
An LLM gateway solves what every AI team reimplements badly: keys, fallback, budgets, caching and observability. On top of that foundation, Jev turns model selection into a 127 ms, $0.00003 classification problem. It's cheap enough to run on every request and accurate enough to trust.
In our scenario, the bill drops from $18k to $6.9k a month. The router costs $32. And with regional inference, the same gateway answers the question legal will ask: where was this data processed?
At Kuaray Tech, we design the AI layer to scale with usage, not with the invoice. Gateways, routing and regional inference are part of our AI and agent engineering work, deployed on the same infrastructure and observability stack as the rest of your platform through our DevOps practice.
Talk to Kuaray about cutting your LLM bill — we measure your real traffic mix, model the routing savings, and put the gateway in place without touching quality.
References
- Vercel — Cost-aware model routing with AI Gateway. https://vercel.com/kb/guide/cost-aware-model-routing-with-ai-gateway
- Arize — What Is an LLM Gateway? https://arize.com/glossary/llm-gateway/
- Vercel Changelog — TypeSafe AI's Jev now available on AI Gateway (Sept 16, 2026). https://vercel.com/changelog/typesafe-ai-jev-now-available-on-ai-gateway
- OpenRouter — What Is Jev? TypeSafe's Decision Model Explained for Developers. https://openrouter.ai/blog/insights/what-is-jev/
- MindStudio — Jev pricing: cost per token. https://www.mindstudio.ai/blog/jev-pricing-cost-per-token
- MarkTechPost — A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out (Sept 23, 2026). https://www.marktechpost.com/2026/09/23/a-coding-guide-to-typesafe-ai-jev/
- Layer3 Labs — Is Jev worth it? (Sept 22, 2026; includes TypeSafe dashboard accuracy and early-access terms). https://www.layer3labs.io/guides/is-jev-worth-it
- Flavio Copes — How much does Jev cost? https://flaviocopes.com/jev-pricing.md
- LiteLLM Docs — TypeSafe AI (Jev) pass-through. https://docs.litellm.ai/docs/pass_through/typesafe
- LiteLLM Blog — JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost (Sept 20, 2026). https://docs.litellm.ai/blog/jev-auto-router-benchmark
- Cloudflare AI Docs — Jev (typesafe). https://developers.cloudflare.com/ai/models/typesafe/jev/
- Classmethod — I tried replacing model routing with TypeSafe (Jev)… (Sept 17, 2026). https://dev.classmethod.jp/en/articles/jev-for-llm-model-routing/
- Anthropic — Claude pricing (accessed Oct 1, 2026). https://platform.claude.com/docs/en/about-claude/pricing
- Vercel Docs — AI Gateway Regional Inference (updated Sept 10, 2026). https://vercel.com/docs/ai-gateway/security-and-compliance/regional-inference
- Anthropic — Data residency (accessed Oct 1, 2026). https://platform.claude.com/docs/en/manage-claude/data-residency
- Cloudmagazin — AWS Bedrock, Anthropic API or self-hosted AI inference (Apr 22, 2026). https://www.cloudmagazin.com/en/2026/04/22/aws-bedrock-anthropic-api-or-self-hosted-ai-inference/
- Vercel Changelog — Regional inference now available on AI Gateway (Jul 27, 2026). https://vercel.com/changelog/regional-inference-now-available-on-ai-gateway
- Vercel AI Gateway —
/v1/modelscatalog (regions and regional pricing, accessed Oct 1, 2026). https://ai-gateway.vercel.sh/v1/models - LiteLLM Docs — Region-based routing (deprecated; use tag-based routing). https://docs.litellm.ai/docs/proxy/customer_routing
- EUrouter — LiteLLM integration. https://www.eurouter.ai/integrations/litellm
Prices and benchmarks checked at the time of writing. Models and prices change fast: rerun the math with current numbers.