How Much Context Does a Decision Model Need?
SEP 29, 2026
How much information does JEV need to make a useful decision, and when is a custom classifier worth training? An exploration of context, labels, and evaluation.

I've been trying JEV this week, and one question kept coming back: How much information does a general-purpose decision model need for a specific task?
The next question matters just as much. If I can give it enough context, what do I gain compared with training a custom classifier? The answer changes when I do not yet have labeled cases or the infrastructure to maintain one.
What JEV is
JEV is TypeSafe's model for structured decisions. The company calls it a System One Model: it takes a state and questions about that state, then returns typed answers. A question can ask it to choose among options, assign a score, or return a bounded value. I use “decision model” to describe the job I give it, not as a claim that this is a universal technical category.
A large language model (LLM) generates text token by token. It can reason about a case, and we can ask it to output a classification in a fixed format. JEV starts with a different interface: the choices and criteria are defined before the request, and independent questions about the same state can be evaluated in parallel. This matters when the answer drives a workflow. A label or score can feed the next step directly; prose requires deciding which part is the answer.
None of this establishes that JEV is more accurate than an LLM, or that a structured answer captures everything important. The comparison I want is what information each needs to get a decision right, the cost of that answer, and how each handles ambiguity. A tidy answer can hide a bad decision; a long explanation can sound convincing without completing the task.
An example separate from my experiment
Imagine a hypothetical support request: "I can't log in." We ask JEV whether it is urgent and which team should handle it. The message alone does not tell us whether one account is affected, the whole service is down, or someone faces a deadline. It also does not define what "urgent" means for this workflow.
I could add the incident's scope, when it started, system signals, and escalation rules. Some of those details might change the decision for good reasons. Others would add noise. An unsupported note saying "this is critical," for example, might sway the answer without adding evidence.
With JEV, some context belongs in the state and some in the question and its criteria. Several independent questions can evaluate the same state in one request. The animation below compares JSON generated token by token with four independent questions: urgency, team, scope, and whether there is enough evidence to start diagnosis. Switch between the message alone and a state with scope and technical signals, while the criteria remain fixed. Results and probabilities are invented to explain the interface, not obtained from model calls; the bars do not measure speed or accuracy.
One state, two interfaces
The same case. Two ways to answer.
Hypothetical case · not a JEV result
"I cannot access the service."
200 accounts affected since the last deployment. The authentication monitor returns 500 errors. No logs attached.
Generative LLM
Sequential textOne output, token after token
{
"urgency": "high",
"team": "engineering",
"scope": "service",
"evidence": "sufficient"
}Complete JSON · ready to consume
It can also produce structured JSON. Values are serialized as a sequence of tokens.
JEV
Independent questionshigh .90 · normal .08 · unknown .02
eng. .92 · support .06 · unknown .02
service .94 · account .04 · unknown .02
sufficient .85 · insufficient .15
Four questions about the same state. None uses another answer; code composes what happens next.
Editorial simulation: invented answers and probabilities, with no model calls or benchmark data. Motion illustrates sequential text versus independent questions, not relative speed or accuracy. Changing context reveals missing evidence; it does not prove that adding fields improves a model.
Given scope and a technical signal, the simulation returns high urgency, engineering, service-wide scope, and enough evidence to start diagnosis. It does not establish the cause.
The interface gives us structured answers. We still have to find out whether the available information supports them. If two questions depend on each other—for instance, choosing a team only after confirming a widespread incident—the workflow must express that dependency. Parallel evaluation does not resolve it by itself.
Finding the context that matters
I would look for sufficient context, not the largest possible state: facts that change an answer for defensible reasons, and missing facts that should trigger abstention or review. Adding context does not necessarily improve a decision when facts, assumptions, and someone else's conclusions are mixed together.
To test what helps, I would hold a set of independently assessed cases constant. I would vary the information in layers: the message alone, then scope, then system signals, then escalation criteria. I would record regressions alongside corrections. Removing fields would show whether something apparently essential actually matters, or whether the model relied on an irrelevant detail.
I would also separate the case state from the question's criteria. The state should contain verifiable observations: how many people are affected, when it began, which services fail. The criteria define “urgent” and the available routing options. If I change both at once and the result improves, I will not know why. If the rules vary between cases, I may end up testing how I phrased them instead of the model's ability.
There is a practical limit. Some cases cannot be decided responsibly from the description and signals available. The workflow should allow “insufficient information” or a review path, and measure how many cases remain unresolved. Forcing a class every time can make completion look better while reducing the system's usefulness.
That separates two questions often treated as one: whether the model understood the task we intended, and whether its answers agree with decisions we can defend. A plausible answer or high confidence does not establish the latter on its own.
The advantage appears before the dataset exists
A specialized classifier would be an attractive alternative if I had enough labeled examples, a stable target definition, and the means to train, evaluate, and maintain it. I could adapt it to the domain and measure performance on cases kept out of training.
But those labels and that infrastructure are often exactly what is missing. JEV lets me try a decision without first training a different model for each question. It may help reveal ambiguous criteria, missing information, and cases that need human review. For a new or low-volume workflow, that can be more useful than starting with a training pipeline.
The choice need not be permanent. I could begin with JEV to explore the task, carefully record decisions and corrections, and revisit a custom classifier once I have enough representative cases. Those records do not automatically become reliable labels. They need shared criteria, review of disagreements, and a test set that is kept separate from tuning.
It cannot turn missing evidence into knowledge. A general-purpose model may fail on local rules or facts it never received. A custom classifier would also struggle with inconsistent labels or incomplete inputs. Choosing between them requires understanding the task, available data, and cost of mistakes.
How I would tell whether it works
I would compare a simple rule as a baseline, JEV with stable state and question definitions, and a custom classifier if I eventually collect enough data. They would face the same held-out cases, set aside before tuning any approach.
I would inspect errors by class, false positives and negatives according to their consequences, stability under irrelevant changes, behavior with incomplete information, latency, and total cost. If I used JEV's confidence to route cases for review, I would need labeled examples to check whether that threshold actually catches errors.
Jevons paradox as a question
TypeSafe says it named JEV after William Stanley Jevons. In The Coal Question, Jevons argued that using coal more efficiently could expand its uses and increase total consumption. That is the intuition behind Jevons paradox: when an activity gets cheaper, additional demand can offset the savings per unit. It is not a universal law, and I am not claiming this has already happened with JEV.
The connection to decision models is useful to me. If each decision becomes cheap and fast enough, we might move beyond automating only high-volume tasks and consult a model at every small step: prioritize a ticket, decide whether to request another detail, choose who handles it next. A cheaper decision could mean many more decisions. Savings per request and the total cost of the system would tell different stories.
That adds another evaluation alongside accuracy: how many new requests appear when friction falls, how many lead to useful actions, how much review work they create, and what resources the whole workflow consumes. The paradox helps frame those questions. Answering them requires real usage data, not an economic analogy.
The right question
It reminds me of a scene from the film I, Robot. Detective Spooner questions a projection of Dr. Lanning, which warns that its responses are limited and that he must ask the right questions. When Spooner finally asks one that opens a path forward, Lanning replies, "That, detective, is the right question."
For me, the right question is not yet "What did JEV answer?" It is "What evidence would show that it answered well, and when would training something specific be worth it?"