From Generative Judges to Decision Models: Does Jev Change the Game for Agentic Measurement and Learning?
Can better measurement of an episode lead to better learning from it?
In our earlier note, “From Agent Evaluation to Episode Learning,” we argued that long-running agents create a different evaluation problem. The relevant object is not an individual response or tool call, but an episode: a sequence of decisions and state changes that eventually produces a business outcome.
This raises a second question. How should we measure an episode while it is unfolding?
The limits of generative judges
A common approach is to use an LLM as a judge. The model receives some portion of an agent's trajectory and is asked questions such as: Is the agent making progress? Has the customer's intent changed? Was the previous action useful?
This works naturally for small traces. It becomes less attractive for an episode containing hundreds of interactions distributed across days or weeks. The judge must repeatedly reconstruct the relevant state from an expanding history. Recent work on long-horizon evaluation suggests that both state reconstruction and evaluation become harder as trajectories grow.
Jev, a decision model introduced by TypeSafe AI, suggests a different approach.
Jev does not generate an unrestricted answer. It answers predefined questions using typed outputs—choices, scores, or binary decisions—and returns probabilities associated with those decisions.
For episode measurement, this distinction is useful.
Measure the state as it changes
Consider a customer acquisition episode. At some point, the system has observed the customer's history, an offer, several messages, and the customer's latest response.
Instead of asking a model to summarize the entire episode, we can ask a small set of fixed questions:
The agent then takes an action. Perhaps it discusses pricing, and the customer responds.
The context has changed. We ask the same questions again:
Repeated through an episode, this produces something different from a conventional trace. It produces a probabilistic history of how the estimated state evolved.
That may be a useful measurement primitive for long-running agents.
From sparse outcomes to continuous measurement
The attraction is particularly strong when the true outcome is delayed.
In the example above, conversion may not be observed for two weeks. If conversion is the only measurement available, most of the trajectory has no immediate feedback.
Repeated decision estimates introduce intermediate signals. We can observe that after a particular interaction, estimated intent increased, the likely objection changed, and estimated conversion probability moved from 0.54 to 0.63.
Across many episodes, we can begin to construct a dataset containing:
This is potentially more useful for learning than either raw traces or a final success/failure label.
It also suggests a division of labour between models. Generative models can perform tasks that require an open output space: reasoning, conversation, planning, and action generation. Decision models can repeatedly measure predefined properties of the environment.
The model that acts need not be the model that measures.
Measurement is not causality
There is an important limitation.
Suppose conversion probability increases immediately after an incentive is offered. We cannot conclude that the incentive caused the improvement.
The customer may already have decided to buy. The agent may offer incentives selectively to particular customers. Another interaction may have caused the change. The decision model itself may simply have revised its estimate after receiving additional evidence.
The distinction can be expressed as:
where Y is the eventual business outcome.
Jev may therefore help answer:
What appears to be happening to the episode?
It does not, by itself, answer:
What caused it to happen?
This is where the episodic-learning problem remains open.
Our research direction
At Spyne, we are interested in whether these ideas can be combined into a learning system:
The first research question is whether typed decision models can measure relevant state variables reliably as enterprise episodes evolve. This requires testing calibration, consistency, domain shift, and the choice of variables themselves.
The second is attribution. Given the measured state changes and eventual outcome, which interventions contributed to the result?
The third is improvement. Can evidence accumulated across many episodes identify changes to prompts, timing, incentives, routing, tools, or workflows?
The final step is experimental. A plausible improvement is still only a hypothesis. It should alter future behaviour only when prospective measurement against an appropriate control provides evidence that it improves the actual business outcome.
Does Jev change the game?
It is too early to know.
Its more interesting contribution may not be a particular model architecture, but a different formulation of the measurement problem. Long-running agent systems contain many questions whose answer spaces are known in advance. Those questions may not require generation. They require decisions under uncertainty.
If decision models can answer them cheaply and with well-calibrated probabilities, long-running agent execution becomes easier to instrument as a sequence of measured state changes.
That does not solve episodic learning. Attribution, generalization, and experimental validation remain.
But it may give them better data.
The resulting research question is straightforward: can better measurement of an episode lead to better learning from it?