{"slug": "automata-from-agent-traces-failure-and-next-step-prediction", "title": "Automata from Agent Traces: Failure and Next-Step Prediction", "summary": "Researchers from an undisclosed institution introduced a method that compresses entire trace corpora of LLM-based agents into a single finite-state machine (FSM), enabling next-step and failure prediction. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness, and build in milliseconds. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion.", "body_md": "arXiv:2608.23670v1 Announce Type: new\nAbstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, so they miss the cross-run topology that links next-step and failure prediction. To recover that shared structure, we collapse an entire trace corpus into a single, compact finite-state machine (FSM) that serves as a structural substrate for the otherwise unpredictable behavior of LLM agents. Across twelve public datasets, the FSMs are compact (7-43 states), replay held-out data at >=0.997 fitness with near-identical topology across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context outperforms Agent Workflow Memory on every ground-truth-matched dataset. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor ranks failing runs above passing ones from a partial trace, triggering early stopping well before completion. Behavioral topology thus appears shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring.", "url": "https://wpnews.pro/news/automata-from-agent-traces-failure-and-next-step-prediction", "canonical_source": "https://arxiv.org/abs/2608.23670", "published_at": "2026-08-26 04:00:00+00:00", "updated_at": "2026-08-26 04:15:37.603145+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-safety", "ai-agents"], "entities": ["arXiv", "Agent Workflow Memory"], "alternates": {"html": "https://wpnews.pro/news/automata-from-agent-traces-failure-and-next-step-prediction", "markdown": "https://wpnews.pro/news/automata-from-agent-traces-failure-and-next-step-prediction.md", "text": "https://wpnews.pro/news/automata-from-agent-traces-failure-and-next-step-prediction.txt", "jsonld": "https://wpnews.pro/news/automata-from-agent-traces-failure-and-next-step-prediction.jsonld"}}