← Back to Blog
CostModel RoutingAgentsTutorialArchitecture

LLM Model Routing in Practice: Cut Your LLM Bill with Jev and jev-router

·3 min read

LLM Model Routing in Practice: Cut Your LLM Bill with Jev and jev-router

Sending every user request to a flagship LLM is how most teams' API bills get out of control. The fix the community keeps converging on is model routing: cheaply judge each request's difficulty and type first, send easy requests to a lightweight model, and escalate only the requests that genuinely need the flagship. The routing step alone usually accounts for the largest share of the savings.

Routing comes in two flavors, and Jev can do both — with two official paths:

  1. Business routing — "which pipeline / queue does this request belong to, what should the agent do next." Use the Jev Decisions API: send state + questions and branch on the result.
  2. Model routing — "which model and how much reasoning effort does this request need." OpenRouter packaged this as typesafe/jev-router: change the model field and you're done.

TL;DR: for deterministic business branching (auditable, threshold-controlled), call Jev yourself. To cut the bill on an existing app without touching business logic, switch to typesafe/jev-router.


Pattern 1: Write the routing logic yourself (Decisions API)

The idea: insert a cheap decision call before the expensive one, have Jev answer a question you define, and branch in code.

For a support agent, start by judging request complexity:

curl https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "typesafe/jev-1.13",
    "state": "User message: I changed my password and still can't log in, tried three devices. Screenshot attached: error code ERR_921.",
    "questions": [
      {
        "type": "score",
        "question": "How much reasoning power does this request need?",
        "levels": ["trivial", "routine", "complex", "critical"]
      },
      {
        "type": "choice",
        "question": "Who should handle this?",
        "options": ["lightweight model", "flagship model", "human agent"]
      }
    ]
  }'

Read selected + confidence + the per-option probability distribution, then:

  • "routine" and below → lightweight model;
  • "complex" → flagship model;
  • "critical", or confidence below a threshold (say 0.7) → a human. Don't guess.

This is the confidence-based tiered routing the community keeps discussing: low confidence escalates automatically instead of paying flagship prices for everything. It works because the decision layer itself is cheap and fast — which is exactly Jev's positioning (input-token pricing, free output, ~70–500ms responses).

The same pattern routes agent actions: feed a DOM snapshot or tool-call result as the state, ask "which tool should be called next", and stop paying GPT to think about every single step.

Pattern 2: The managed router, typesafe/jev-router

If your scenario is "I don't want to design questions myself, I just want every request served by the cheapest sufficient model," OpenRouter ships this as a ready-made router: typesafe/jev-router.

How it works: Jev reads the conversation, judges the task type, difficulty, and how much a stronger model would help, then picks the cheapest model in your candidate pool that meets the bar; unusually hard requests get an "expert advisor" model as backup. It speaks standard Chat Completions, streaming included:

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-OpenRouter-Metadata: enabled" \
  -d '{
    "model": "typesafe/jev-router",
    "messages": [
      { "role": "user", "content": "Categorize the errors in this log excerpt" }
    ]
  }'

The response's model field tells you which model actually served the request. With X-OpenRouter-Metadata: enabled you also get routing details like resolved_models and list_fallback for debugging.

Controlling the candidate pool (optional jev-router plugin fields):

Field Effect
models / allowed_models Include list — only matching models can be selected
excluded_models Exclusion list — never selected, even if also in the include list

Three engineering semantics to know:

  • Include lists accept exact slugs, dated revisions, anthropic/* wildcards, and ~author/family-latest aliases, up to 1,024 patterns per list;
  • An include list that matches nothing doesn't error — the router falls back to the default pool (reported as list_fallback: "models_ignored" in metadata);
  • Exclusions always apply — exclude every pool model and the request fails with a 404.

Which pattern to pick

Call Jev, write your own routing typesafe/jev-router
The question it answers Business: which queue, which pipeline, what next Model: which LLM serves this request
Integration cost Define state and questions, write the branching Change one model string, business code untouched
Control Fully custom options, thresholds, fallback logic Include/exclude lists + the official selection policy
Typical use Ticket triage, intent detection, agent action decisions Cost-cutting for existing chat apps and agent frameworks

The two stack: use the Decisions API for deterministic business branching, and jev-router to cut the model bill underneath.

The math

All figures below are illustrative (live prices live on the OpenRouter model pages), but the arithmetic is general. Assume: flagship $3 / 1M input tokens, lightweight $0.1 / 1M, Jev $0.042 / 1M; each request is 500 tokens, each routing decision 300 tokens, volume 1M requests/month.

  • All flagship: 1M × 500 ÷ 1M × $3 = $1,500
  • Routed (70% lightweight / 30% flagship): 350M × $0.1 + 150M × $3 = $35 + $450 = $485, plus Jev's routing judgments at 300M × $0.042 ≈ $12.6 → about $498
  • A ~67% reduction — and that ignores output tokens, where Jev is free and LLMs bill per token.

Your actual gain depends on the share of "downgradable" requests. Support, tagging, formatting, and simple Q&A traffic usually ranks high; genuinely hard tasks (complex reasoning, long-form writing) are beyond what a routing layer can help with — which is exactly why you set confidence thresholds.

Engineering caveats

  • 32k token context cap (state + questions combined). Summarize or truncate long conversations before routing.
  • Jev doesn't explain its answers. When you need the "why," have code call a chat model with the structured result, or escalate low-confidence cases to a human.
  • Versioning: pin typesafe/jev-1.13 in production for reproducibility; use the ~typesafe/jev-latest alias to auto-track new releases.
  • The routing layer is also a failure point: give Jev / jev-router calls a timeout and a fallback path (on timeout → default to a model you trust). Don't let the cost-saving layer become the outage.

Further reading