Article

Jev: Zero-Shot Classification Without the Text Generation

Author
Josh Fjelstul, PhD
Category
Analysis
Date
October 6, 2026
Time
8 min

A lot of people use LLMs and zero-shot classification to make decisions, but that takes a lot of compute, and it can be really slow when you're working with a large corpus. Let's take a look a Jev — a new zero-shot classification model that's cheaper and faster than using an LLM.

LLMs are increasingly used for making small, well-scoped decisions in products — like routing or tagging. But using an LLM is a roundabout way of doing that. If you're trying to get an ML model to answer a well-defined decision-theoretic problem, you'll have a few elements: context, a question, and a set of pre-defined choices. With LLMs, you have to write a prompt, run inference, parse the output, validate the output, and then do something with that output. It's expensive — you have to pay for input tokens and output tokens, even though you don't actually need any text generated. You just need the model to choose the best option from the choice set.

A new approach

Let's take a look at a new model called Jev, which tries to solve the problem of slow, wasteful text-generation when using LLMs to make well-defined decisions. Jev is a proprietary model from TypeSafe AI.
Jev is a zero-shot classification model. The goal of zero-shot classification is to assign labels without training on examples for your specific task. Jev models a probability distribution over a set of answer choices as a function of the state and a question. The input is a JSON document that includes the state, a set of questions, and a set of options for choice questions. The state can be a string or JSON.
Since Jev is a proprietary model, we don't know exactly what the architecture is. We know it's fast and low-cost. And we know you can process questions in parallel. TypeSafe claims it built a new model architecture. But you could build something similar using an open-weight LLM. You could run a forward pass over the state, question, and options and produce a probability distribution over the options.
Jev supports three kinds of questions:
  • Choice questions: Returns the selected option, a probability distribution across all options, and a confidence value.
  • Score questions: Returns a probability for each rubric level, a probability-weighted expected score, and a confidence value.
  • Yes/No questions: Returns a single probability (that the answer is "yes"). TypeSafe calls these "noul" questions.
Choice questions can have up to 255 options, score rubrics take 2-10 levels, and the state plus the longest question can be up to 32k tokens.
Note that a confidence value isn't the same thing as probability. It's a summary of how tight a probability distribution is. For choice questions, TypeSafe rescales the top probability so that a uniform distribution gets a confidence of 0 and a certain answer gets a confidence of 1. With two options, a top probability of 0.75 yields a confidence of 0.5. If you want to know how likely it is that the chosen option is correct, you have to look at the probability, not the confidence score.
Jev is cheaper and faster than zero-shot classification with an LLM. It's $0.042 per million input tokens, and there's no charge for output. TypeSafe claims 70–500ms response times.

How to think about it

Jev is basically zero-shot classification, just without any text generation.
If you've worked with BERT models, the closest analogy is a multiple-choice head — not a standard multi-class classification head. A standard classification head has a final layer with a fixed number of output nodes — one per available label. If you want to add a label, you have to retrain the model.
When you use a multiple-choice head, BERT pairs the context with each candidate answer, gives each pair a score, and returns a probability distribution over the candidates. Because the candidates are passed in as input text, the same head can handle new choices or a different number of them.
The difference is that a BERT multiple-choice model is fine-tuned for a specific task (and may not transfer well to other tasks), whereas Jev doesn't need to be fine-tuned: you just describe the question and the options in the input, and it scores them without any task-specific training.

What's new about it

What's new about Jev is the training objective. TypeSafe says it uses a reinforcement learning method called reinforcement learning for calibrated decisions (RLCD), which optimizes for well-calibrated probabilities, but it hasn't published the reward function, training data, or optimization method.
A model is well-calibrated when its predicted probabilities are consistent with how often things actually happen. If a model says the probability of an outcome is 0.8, we should expect that outcome to occur 80% of the time.
A model can be highly accurate and badly calibrated. Many ML models have good accuracy — they get their predictions right most of the time — but they report high confidence even when they get something wrong. An overconfident model will put too much probability mass on the most likely outcome — and an underconfident model too little.
The flip side to calibration is sharpness, or discrimination. A model that always predicts 30% for something that happens 30% of the time is well-calibrated, but it doesn't tell you anything useful about any particular situation.
So, an accurate model isn't necessarily well-calibrated and a well-calibrated model isn't necessarily accurate. Ideally, you'd like a model that's well-calibrated and discriminatory.
The goal of RLCD is to reward calibration.
We can compare RLCD to several other types of reinforcement learning: reinforcement learning from human feedback (RLHF), reinforcement learning on verifiable rewards (RLVR), and reinforcement learning with calibration rewards (RLCR). The goal of RLHF is to fine-tune a model to generate answers people prefer. This often makes calibration worse because people prefer confident-sounding answers. The goal of RLVR is to fine-tune a model to generate correct predictions. The idea is to use a reward function that rewards correct answers and punishes incorrect answers. The goal of RLCR is to fine-tune a model to generate predictions that are correct and well-calibrated — a similar goal to RLCD. The idea is to change the reward function so that it also rewards calibration.
So, RLHF rewards approval, RLVR rewards correctness, and RLCR rewards correctness and honesty. RLCD is closest to RLCR. The main difference between RLCD and RLCR is where the calibration happens. RLCR calibrates a confidence score in an LLM's text output. Jev appears to calibrate its output distribution over typed answers.
The question is: Why should we care about calibration? This comes down to how you're using the model. If a false positive or false negative is particularly costly, you might want to know how confident the model is before acting on its predictions.
One natural way to train for calibration is to use a proper scoring rule. A scoring rule is any function that scores how good a prediction is. A proper scoring rule is a function where the reward is maximized when a model accurately reports its true belief, given the data — which is different from just being accurate. Under a proper scoring rule, a model that inflates its confidence or hedges its predictions receives a lower expected reward. We don't know if TypeSafe's RLCD uses a proper scoring rule, but that would be a natural way to implement RLCD's stated objective.

How well it works

How well Jev works — compared to the alternatives — is still an open question. But we do have some evidence. For example, in an independent evaluation, Jev scored 95–99% accuracy on standard benchmarks like IMDB, SST-2, HellaSwag, and ARC. But results varied widely across other tasks. It beat an open-weight Qwen model, but the margins were small.
Like other models, it struggled with fine-grained labels, low-resource languages, and decisions that required domain expertise. For choice questions, the probabilities were well-calibrated, but for yes/no questions, the threshold needed to be tuned. Other evaluations have found mixed results.

Why you might want to use it

There are two main benefits of using a model like Jev.
First, it's cheaper and faster. With LLMs, you have to run inference to generate the tokens for the JSON output, which adds latency and cost. Jev doesn't generate text at all — just predicted probabilities. You don't have to generate any JSON output — so you don't have to pay for those output tokens.
Second, you get calibrated probabilities (or that's the claim, at least), not just a label. When you use an LLM to do zero-shot classification to make a decision, you don't know how confident the model is. Some LLM APIs return token logprobs, but logprobs from RLHF-tuned models can be poorly calibrated. LLMs can't self-report their own confidence in their own output reliably, unless they've been specifically trained to do that. If some options in the choice set are more risky or more costly than others, you might want to know how confident the model is before acting on its decisions.
You might see people talking about the structured output being an advantage, but it really isn't. When you're doing zero-shot classification with an LLM, one of the most common failure modes is that the LLM's output doesn't match the schema you provided. That's why it's important to validate the LLM's output to make sure it conforms to your schema. If you don't validate the output, you run the risk of the LLM producing data that causes an error when you try to use it downstream, like when you display it in an app. But most LLM APIs offer structured outputs or constrained decoding, which largely solves this problem.

What the limitations are

You might see people say that Jev's biggest limitation is that you have to have a well-defined decision-theoretic problem: you need relevant context, a clear question, and pre-defined choices. This might seem like a lot more work than using an LLM for zero-shot classification, but it's really not. If you're doing zero-shot classification well, you'll already have a carefully-designed schema and clearly-defined labels. You're just giving the model this information as input, rather than trying to force the model to conform to the schema.
There are other real limitations. Jev works best when a decision can be correctly made based on a short description of the options. It doesn't do as well on fine-grained labels and labels that are poorly conceptualized. Like any classification problem, measurement strategy is critical — if your choices aren't clearly defined and clearly distinct from each other, the model will struggle. Jev also struggles with tasks that require multi-step reasoning.
In short, Jev is useful for applications where (1) you need to make quick, simple, real-time decisions between a limited set of clearly-defined options, or (2) you're working with a large corpus where doing zero-shot classification with an LLM would be too expensive.

Takeaways

A lot of people use LLMs and zero-shot classification to make decisions, but that takes a lot of compute, and it can be really slow when you're working with a large corpus. LLMs generate a lot of tokens — especially reasoning models — and if you're doing zero-shot classification, you don't need any intermediary text.
RLHF-tuned models can produce token probabilities that are miscalibrated for zero-shot classification tasks because they're specifically fine-tuned to generate answers people like, and people like confident answers. RLCD — however it works — is designed to produce a model that's better-calibrated, which can help you make better decisions by giving you the information you need (well-calibrated probabilities) to manage risks.
So, for some applications, Jev is faster, cheaper, and more informative. But you still have to do the hard work up-front — you have to have a well-specified decision-theoretic problem with a well-defined choice set. Getting that right is the hardest thing — and the most important.
Recommended posts
Ready to build?

From Problem Framing to Production.

Whether you need a domain-adapted text classification model, or an end-to-end recommender system with RAG, I help you ask the right questions, frame your business problem, and build cutting-edge AI/ML solutions.
Schedule a consultation