How Jev Cut Our Classification Cost by 20×

Max Heckel - Author profile picture
Max Heckel
· 7 min read
EngineeringAI agentsTypeScriptClassification

We moved classification calls inside Ari from general-purpose LLMs to TypeSafe's Jev. For the workloads we've moved, we're seeing roughly one-twentieth the cost and 10–15× faster decisions. That's about a 95% reduction in classification cost.

We had been paying for LLM calls to route messages, detect topic changes, decide whether a follow-up should interrupt ongoing work, and check whether a response satisfied the request. An agent can make several of these decisions before the user gets an answer.

Those savings apply to the classification calls we replaced. The agent still spends time and money executing tools, generating text, and running fallback calls. Our implementation keeps those paths and lets Jev handle decisions with a fixed set of possible answers.

Langfuse chart showing average IntentDetection latency at roughly 1.4–1.5 seconds and JevIntentDetection latency at about 0.2 seconds on September 28.

Intent detection in Langfuse: average trace latency drops from roughly 1.4–1.5 seconds for IntentDetection to about 0.2 seconds for JevIntentDetection in this window.

Give a decision its own interface

We call TypeSafe's evaluation endpoint directly. The request contains a model, a state holding the evidence, and a map of typed questions. Answers come back under the same question names. We use two question types: choice, for choosing among defined options, and noul, for a yes/no probability. See the TypeSafe API reference for the wire contract.

We supply the evidence separately from the question instructions, then interpret the answers in application code.

For intent detection, one request asks five questions:

QuestionAnswer
What kind of request is this?default, plan, still_working, or simple
Does it start a new topic?Yes/no probability
Did the user explicitly request a threaded reply?Yes/no probability
Has the user's conversational need changed?A choice including unchanged
Do we need a new conversational hint?Yes/no probability

The questions share the conversation evidence. Each has focused instructions, so a decision about threading doesn't have to compete with instructions to write a friendly acknowledgement.

“Thanks” while waiting for work belongs on a different route from “thanks” after the work is done. “And in Austin?” can extend a weather request even though it doesn't repeat the word “weather.” We spell out those distinctions in the prompts.

Carry the question types into the answers

TypeScript can derive the answer type from the question map, keeping each answer tied to the options its question allows. Here's an example of that pattern:

type Question = | { type: 'choice'; criteria: Record<string, string> } | { type: 'noul' }; type Answers<Q extends Record<string, Question>> = { [K in keyof Q]: Q[K] extends { type: 'choice'; criteria: infer C } ? { type: 'choice'; choice: keyof C & string; confidence: number; probabilities: Record<keyof C & string, number>; } : { type: 'noul'; noul: number }; };

Defining questions with as const preserves their literal keys. That means an intent answer's choice can be the union of the four routes declared in the question. Callers get the corresponding probability keys and answer fields without maintaining a separate response interface by hand.

A request function can return Promise<Answers<Q> | undefined>, making the distinction between a valid answer and an unavailable decision explicit when choosing a fallback.

TypeScript alone doesn't validate an HTTP response. Runtime validation with Zod and additional distribution checks ensures every requested answer is present with the right discriminant. Probabilities must be finite numbers between zero and one. A choice must belong to the declared options, its distribution must contain exactly those options, and the probabilities must sum to one within a small tolerance. The selected option must have the highest probability, allowing ties within tolerance.

Malformed responses take the unavailable path. A response that passes every check can still contain a wrong classification.

Uncertainty needs a policy

A probability-to-boolean conversion can deliberately leave room for uncertainty:

function classifyBoolean(value: number): boolean | undefined { if (value >= 0.9) return true; if (value <= 0.1) return false; return undefined; }

A result in the middle is uncertain. We handle it differently depending on the decision.

For intent routing, confidence below 0.9 selects the normal default route. Uncertain topic and threading decisions become false. An uncertain conversational-style change leaves the existing style alone. These are conservative routing defaults; low confidence by itself doesn't trigger another LLM classification call.

An unavailable Jev response does trigger the existing LLM classifier. A timeout, invalid response, disabled integration, or missing configuration is different from a valid answer with low confidence.

Interrupt detection has a different policy because overlooking a user's correction can leave the agent doing the wrong work. We accept an interrupt probability of at least 0.8, accept a confident no at 0.1 or below, and send the uncertain middle to the existing LLM path.

These thresholds are rollout policy. A threshold of 0.9 doesn't establish 90% accuracy on our traffic.

Keep text generation where it earns its cost

Our previous intent call combined classification with generated text, including acknowledgements and short hints about the user's conversational needs. Moving the classifier meant separating those responsibilities.

Once Jev returns a decision, an ordinary default route can proceed without an additional intent-text call. Routes such as plan and still_working still request generated text. Certain other non-default decisions can request a hint too.

That smaller text-generation call receives the decisions already made and is instructed to generate only an acknowledgement and a short conversational hint. It doesn't get to change the routing.

The common path finishes classification without also paying for prose. When the interaction needs personalized language, we retain a model call for that job.

Use a fast check before expensive verification

We use the same primitive for response verification. The evidence includes the current request, assistant response, and tool results, plus context appropriate to a conversation or workflow.

Jev answers six focused questions: was the request fulfilled, was a completed action claimed, was a result misrepresented, was execution missing, was the response incomplete, and was a required tool failure left unretried?

The fast path accepts a response only when fulfillment scores at least 0.8, the completed-action question is confidently false, and none of the four failure checks is confidently true. Otherwise, we run the existing verifier.

“I sent the email” needs the fuller verification path even if the response otherwise looks successful. Jev identifies that a claim exists; the fallback can do the claim extraction and verification work.

An uncertain completed-action answer also forces fallback. Uncertainty on one of the four failure flags does not independently force it. Keeping these conditions together makes the entire acceptance rule easier to review.

Make failure bounded and visible

Each classification request has a three-second default timeout covering the request and its retries. Rate-limit and overload responses (429 and 529) are retried with at most three attempts, respecting Retry-After within that overall deadline. Other unsuccessful responses take the unavailable path. Cancellation propagates so cancelled work doesn't quietly start a fallback.

The integration can be disabled independently, and the model is configurable. That gives us a way to roll back or change models without changing the decision policies.

We trace classification calls in Langfuse and Phoenix alongside the rest of the agent's work. A decision trace contains the supplied evidence, questions, and returned answers or unavailable status. Credentials and transport options stay out of that payload. We can inspect the classifier in the context of the run it affected.

Testing covers malformed answers, unknown choices, invalid distributions, uncertainty thresholds, retries, timeouts, cancellation, and fallback behavior. Those checks protect the integration contract and routing policy. Classification quality still needs evaluation on representative conversations.

Share this post
Max Heckel - Author profile picture
Max Heckel

Max Heckel is the founding engineer and CTO of Ariso. Before starting Ariso, he worked at Google, McGraw Hill, JupiterOne, and created SciSummary.

LinkedIn

Ready to try Ari?

The AI player-coach that gives every employee the tools to lead themselves.

Try Ari Free