OnTrack: Real-Time Agent Monitoring via Streaming Optimal Transport A developer introduced OnTrack, a real-time agent monitoring system that uses streaming optimal transport to compare an agent's unfolding execution trajectory against reference runs, flagging divergence in roughly a millisecond per step. The approach tracks the dependency graph between steps rather than just tool calls, distinguishing benign reordering from dangerous divergence, and can alert or block execution when divergence exceeds a configurable threshold. It is evaluated on SWE-bench and targets high-frequency and cost-sensitive deployments where post-hoc log analysis is too slow and shadow LLM safeguards add 200-500ms of latency and double inference cost. Autonomous agents execute irreversible actions. A stock trading agent places orders. A trip planner books flights. An IT triage agent restarts production services. Current safeguards are either too slow post-hoc log analysis or too expensive a second LLM watching every step . OnTrack introduces a third option: streaming trajectory comparison that flags divergence in about a millisecond per step, before the damage compounds. The core idea is structure-aware optimal transport. Instead of waiting for task completion or running a shadow agent, OnTrack compares the unfolding trajectory against reference runs in real time. It tracks not just what tools are called, but the dependency graph between steps. When an agent starts looping, stalling, or deviating from known-good patterns, the system raises an alert or blocks execution. Post-hoc observability tells you what went wrong after tokens are burned and APIs are hit. Safeguard agents a second LLM evaluating each step add 200-500ms of latency and double your inference cost. Neither works for high-frequency trading agents or cost-sensitive production deployments. OnTrack targets the gap between action and consequence: The system operates in three regimes based on available data: | Access Level | Available Data | Detection Capability | |---|---|---| | Full | Historical runs + tool schemas | Plan violation, dependency graph mismatch | | Intermediate | Tool schemas only | Loop detection, repeated tool calls | | Minimal | Step logs as generated | Stall detection, basic anomaly flagging | Traditional optimal transport computes the minimum cost to transform one distribution into another. OnTrack adapts this for streaming agent trajectories by comparing partial execution graphs in real time. State representation : Comparison mechanism : The structure-aware component is critical. Two agents might call the same tools in different orders but produce equivalent results. OnTrack's graph comparison distinguishes between benign reordering and dangerous divergence. Simplified trajectory comparison pseudocode class TrajectoryMonitor: def init self, reference graphs, threshold=0.3 : self.references = reference graphs self.threshold = threshold self.current graph = DependencyGraph def observe step self, step : Add step to current trajectory graph self.current graph.add node step Compute streaming OT distance to nearest reference min distance = min optimal transport distance self.current graph, ref for ref in self.references Trigger intervention if divergence exceeds threshold if min distance self.threshold: return InterventionSignal distance=min distance, action="BLOCK" if min distance 0.5 else "ALERT" return None The system needs examples of successful runs. For stock trading agents, this creates a bootstrapping problem: what counts as a "safe" reference when optimal strategies are contested? Three approaches : The paper evaluates OnTrack on SWE-bench, where success is well-defined tests pass . In financial domains, you need explicit policy: is a 2% loss acceptable exploration or a failure to intervene? Blocking an agent mid-execution has consequences. If the agent was exploring a novel but valid solution path, intervention wastes the partial work. If it was heading toward a catastrophic action, blocking saves money and reputation. OnTrack's evaluation on SWE-bench shows the trade-off: For financial agents, you tune the threshold based on risk tolerance: OnTrack runs as a sidecar process or inline middleware in the agent orchestration layer. It does not replace your LLM or tool execution runtime. Integration points : Latency characteristics : For high-frequency trading, this fits within the execution budget. For trip planning or IT triage, it's negligible. When OnTrack blocks an agent, you need to understand why. The system exposes: This turns intervention from a black-box block into a debugging signal. If OnTrack repeatedly flags a new tool usage pattern that succeeds, you update your reference set. Stale references : If your agent's environment changes new APIs, updated tool schemas , old reference trajectories become misleading. Solution: continuous reference set refresh from recent successful runs. Adversarial agents : A sufficiently capable agent could learn to mimic reference trajectory structure while pursuing a different goal. OnTrack is not a security boundary; it's a safety rail. High-variance domains : In creative tasks content generation, exploratory research , there may be no canonical "good" trajectory. OnTrack works best when success has structural invariants. Cold start : With zero reference trajectories, the system can only detect loops and stalls. You need at least a handful of successful runs to enable plan violation detection. Use OnTrack when : Avoid OnTrack when : For financial trading agents, trip planners, and IT triage, OnTrack offers a practical middle ground: real-time intervention without the cost of a shadow LLM. The key is curating reference trajectories that capture your risk boundaries, not just historical success.