All posts
guide

LLM routing with decision models: test before you switch

Give your router clear choices, test its fallbacks and count the whole bill. Compare Decisions API, Jev and a small model before you switch.

EiliyaOctober 1, 20268 min read

A router can choose a handler without writing an essay. Decision models for LLM routing take the relevant state and a closed set of answers, then leave your code to dispatch the work.

OpenAI's Decisions API announcement puts that pattern alongside TypeSafe AI's Jev and the familiar small-model classifier. The right choice depends on the decision you're making. A bounded router can reduce unnecessary generation and make outcomes easier to inspect, but it can't fix unclear categories or guarantee a correct action.

Spend LLM routing calls on ambiguous requests

Start with the result your app needs. An order check might go to a database lookup, a product question to a specialist model, and a complaint with missing context to clarification or human review.

TypeSafe's intent-routing documentation describes this division of work: the router selects a handler, and the handler performs the task. Keep those jobs separate so you can evaluate the choice itself.

For tool selection within an agent, offer the currently relevant actions. Your code supplies the known state and maps the choice to an implementation, keeping the model's job narrower than deciding everything about the next step.

For moderation, a closed answer set might identify cases for review. Evaluate that screening decision against your policy; faster classification alone doesn't establish that a general decision model meets its requirements.

Use code directly when the answer follows from a known rule. Account permissions, arithmetic and available inventory belong there. Save the model call for interpreting language or assessing an ambiguous situation.

Give a strict enum a fair shot first

Your weakest baseline is a large model asked for JSON, followed by parsing and more prompts when the format changes. A decision service may avoid some of that work.

You have a stronger baseline available. OpenAI's Structured Outputs guide documents schema-conforming responses, including enum values, and explicitly describes avoiding retries for incorrectly formatted responses. Your code still needs to handle refusals, and a conforming result can contain a semantic mistake.

Compare decision models with a small model returning a short strict enum too. If your router asks for a long explanation it never uses, remove that requirement before measuring a new provider's advantage.

A specialized decision model gives you an interface built for decisions and a chance to improve efficiency. Measure those benefits on your workload. Bounded output alone doesn't establish an advantage over every generative model with the same constraint.

Decisions API, Jev or a small model?

As of September 30, 2026, OpenAI's recap announces finite-answer Luna decisions with text or image context, in limited preview. A dedicated public request schema and separate price remain unconfirmed.

Jev gives you a documented HTTP contract and typed questions. Its API reference describes Choice, Score and Noul responses. Choice and Score include probabilities and confidence; Noul is a yes/no probability without a separate confidence field.

Approach Reason to evaluate it What still needs checking
OpenAI Decisions API Announced image support and Luna-based decisions Account access, actual contract, price and performance
Jev Documented probability outputs and parallel questions Accuracy, calibration and service behavior on your traffic
Small model with strict output A controllable baseline using an existing model integration Semantic errors, refusals, latency and cost
Deterministic code Exact rules with no semantic ambiguity Whether the rules cover the incoming request

Keep your app's routing contract independent of the provider. You can then compare the same accepted labels, test cases and handlers, changing only the decision adapter.

Design the answer set before writing the prompt

This support-router contract illustrates the design. These are application labels, not a request body for either vendor.

Label Use it when Handler
order_lookup The user requests existing order status Authorized order lookup
product_advice The user asks about product suitability or features Product specialist
billing_review The message concerns charges or invoice disputes Billing workflow
clarify The request lacks information needed to choose Clarification response
human_review The case requires review under application policy Human queue

Define the boundaries between neighboring labels. If a message mentions both a missing delivery and a disputed charge, which takes priority? Without that policy, your reviewers may disagree before the model even sees it.

Keep topic separate from action when they answer different questions. A message can concern billing without authorizing a refund. Recognizing intent shouldn't implicitly approve the requested transaction.

Give unsupported requests a real path. If your router can only choose specialist handlers, it must pick one even when none fits. An explicit clarification or review outcome lets your app handle that case deliberately.

TypeSafe's question guide recommends an other or equivalent answer when the choice set may not cover the input. Decide what your app does with that label, and include labeled examples in your evaluation set.

Supply relevant state and keep permissions in code

Send the evidence your router needs: the current request, relevant prior messages, available handlers and facts that alter the routing policy. Avoid dumping an entire account history into every call.

For tool choice, generate the eligible candidates in code. Remove tools that are unavailable or that the account can't use. Check permissions again when executing the selected handler, because the account state may have changed.

Keep exact comparisons outside the model. TypeSafe's Jev limitation page warns about numerical precision, dates, distracting state and adversarial content. The advice: keep arithmetic in code and make criteria explicit.

Ask the model what the message means, then use trusted records to decide what your app permits. Malicious input can still push a model toward the wrong allowed label, even with a closed list.

Test what confidence tells you about errors

Jev's confidence documentation distinguishes a probability distribution from the confidence statistic derived from it. You shouldn't read every confidence value as a measured probability that the chosen label is correct.

Build your policy around observed errors. Allow automatic routing where tests justify it; escalate ambiguous cases to a stronger model, ask for missing context or send them for review. A reversible queue assignment and an action with financial consequences need different policies.

The following is provider-neutral pseudocode, not an SDK example:

eligible = allowed_handlers(account, workflow_state)
decision = evaluate_route(relevant_context, eligible)
if decision_failed_or_incomplete(decision):
    return fallback_for_service_failure()
if review_required(decision, tested_policy):
    return review_or_clarify()
return execute_with_permission_checks(decision.label)

Keep service failure separate from uncertainty. A timeout tells you nothing about which category the message belongs in. If a provider doesn't expose confidence, use explicit review labels and test their behavior instead of inventing a numeric score in your adapter.

Count fallbacks, slow requests and downstream bills

Create a labeled set before tuning the option descriptions. Include clean cases, overlapping intents, missing context, unusual wording and hostile instructions inside the content. Reserve a held-out set for the final comparison, so prompt changes don't quietly become training on the test.

Use these measures together:

  • Routing accuracy, with results broken down by label.
  • Wrong automatic actions and how serious they are.
  • Coverage: how much traffic proceeds without fallback or review.
  • End-to-end latency, including slower requests and fallbacks.
  • Billed cost for the complete workflow, including downstream calls.

When you have probabilities, compare their ranges with observed outcomes. Check whether confidence tracks accuracy for each important route. One overall accuracy figure can hide a poor billing route behind many easy order queries.

Measure cold calls and realistic concurrency from the region where your app runs, including preprocessing. If a text-only router needs another model to describe an image, count that work when comparing it with a router that accepts images directly.

Add router calls, fallback calls and downstream work to get the full cost. A router earns its place when that total improves while errors stay acceptable. The input token price alone won't tell you that.

A recipe router shows a boundary worth testing

On Reddit, developer blackbarata described a Jev router for a personal knowledge app. It chose a scraper when a recipe transcript lacked full instructions and a recipe agent when the instructions were present.

The author explicitly called it two examples rather than an evaluation. They show a plausible decision boundary, without establishing reliability under load or across ambiguous requests.

Use reports like this to find cases worth testing. Replay your own traffic in shadow mode, compare the new router with your existing path and inspect disagreements before letting it control production actions.

Fund downstream calls with a user wallet

For downstream text tasks, AI Pass connects Claude, GPT and Gemini through an OpenAI-compatible API at https://aipass.one/v1. Its agent tools use one account and balance without provider keys. Users fund exact usage through a portable wallet, with no required subscription. Keep your decision adapter separate; AI Pass is independent and unaffiliated with OpenAI.

With Sign in with ChatGPT, the user's plan covers OpenAI usage without revenue for you; AI Pass lets you add a markup and earn on eligible paid usage under the current AI Pass terms. Start today at the docs: paste one brief into your coding agent, let it inspect and set up the project, then approve in your browser, without an application, waitlist or manual OAuth client setup.

FAQ

What is LLM routing?

It's choosing which model, tool or handler should process a request. Your router can also choose deterministic code, clarification or human review.

Are decision models always faster than a small LLM?

There's no established universal comparison. Test a small model with short, strict output against the decision service, using the same task and the complete request path.

Can a decision model choose an agent's next tool?

Yes, when the options are defined and the relevant state is supplied. Code must still validate permissions, arguments and execution conditions.

Should moderation rely on confidence alone?

No. Evaluate policy-specific errors and escalation behavior. A confident label is not a substitute for testing the moderation decision on representative content.

Sources