From Agent Evaluation to Episode Learning
How long-running enterprise episodes can provide evidence to discover better policies—and demonstrate that they work.
The Research Problem
Most agent evaluation is local. Given a state, did the agent produce the correct response, select the appropriate tool, or take a reasonable action?
Enterprise work is rarely local.
Consider a financial institution attempting to sell an additional lending product to an existing customer. An episode may begin with customer and product data, continue through messages and conversations over several weeks, transition between autonomous agents and human representatives, introduce incentives, encounter objections, and eventually terminate in conversion or abandonment.
We can represent such an episode as a trajectory:
where G is the objective, st the state at time t, at an action, ot an observation, and Y the terminal business outcome—conversion in this example.
The evaluation problem is no longer simply whether each at was appropriate given st. It is whether the trajectory constituted an effective strategy for achieving G.
This distinction is consequential. An episode can succeed despite poor intermediate decisions. Conversely, every individual action may appear reasonable while the overall episode fails.
Local correctness does not imply global effectiveness.
From Outputs to Trajectories
Recent research is beginning to expose different dimensions of this problem.
AMA-Bench shows that long-horizon agents must preserve objective state and causal relationships that similarity-based memory systems frequently lose. HORIZON examines where failures occur as task horizons increase. Plan-RewardBench demonstrates that evaluators themselves become less reliable as trajectories become longer. Work on agent drift examines how intent, behavior, and coordination change over time, while recent trajectory-attribution research asks which components—and chains of components—contributed to an eventual outcome.
Together, these results suggest that evaluating long-running work cannot be reduced to providing a larger trace to a more capable LLM judge.
The harder problem is learning from the trajectory.
Why This Is Hard
Suppose a customer journey contains the sequence:
What caused the conversion?
Perhaps it was the incentive. Perhaps the human conversation. Perhaps the first message established intent and everything afterwards was incidental. Perhaps the customer would have converted without any intervention.
The observed trajectory cannot answer this because the most important trajectory is unobserved:
What would have happened had the system acted differently?
This creates several fundamental challenges.
Partial observability. Important state variables—customer intent, urgency, or price sensitivity—are latent and must be inferred from imperfect observations.
Sparse and delayed reward. The business outcome may arrive days or weeks after the decisions that influenced it.
Confounding. Actions are not randomly assigned. An agent may offer incentives precisely to customers it predicts are unlikely to convert. Historical correlation can therefore produce misleading conclusions about whether the incentive was effective.
Long-range credit assignment. Outcomes may result from combinations of decisions separated by many intermediate events. The relevant unit of attribution may therefore be a sequence of decisions rather than a single action.
These are not simply evaluation problems. They are problems of causal inference and sequential decision-making.
Our Hypothesis
Our hypothesis at Spyne is that long-running enterprise episodes can become a substrate for learning if we can reconstruct state, attribute outcomes to interventions, generate candidate improvements, and experimentally establish whether those improvements work.
We are investigating four related directions.
First, episode reconstruction. Enterprise trajectories are fragmented across agent traces, conversations, data warehouses, CRM systems, human actions, and downstream transaction systems. We are exploring representations that preserve not only temporal order but relationships between goals, states, decisions, actions, and outcomes.
Second, trajectory attribution. Given an outcome Y, can we identify the decisions that materially contributed to it? This requires moving beyond temporal correlation toward counterfactual reasoning, matched episodes, controlled interventions, and other approaches to estimating contribution.
Third, policy improvement. If repeated episodes reveal a pattern, can the system translate that evidence into an actionable modification—a different prompt, timing rule, incentive strategy, routing decision, tool choice, or workflow?
A related question is when an experience should generalize. An intervention that works in one episode should not automatically become a universal rule. Learning requires identifying the state and context under which an experience remains applicable.
Finally, measurement. A proposed improvement cannot be validated using the same episodes from which it was inferred. Candidate policies must be evaluated prospectively against appropriate controls. Where possible, randomized interventions and A/B experiments provide stronger evidence that a policy change caused an improvement rather than merely correlating with one.
From Evaluation to Learning
This suggests a broader loop:
Suppose the existing system generates episodes under policy π, and Y denotes the terminal business outcome. The objective is ultimately to discover an improved policy π′ such that:
while satisfying the business, operational, and policy constraints governing the episode.
The central problem is therefore not evaluation alone. It is learning under delayed outcomes from partially observed, confounded trajectories.
Our research asks whether long-running enterprise episodes can provide sufficient evidence to discover better policies—and, critically, whether those improvements can be demonstrated rather than merely judged.