{"slug": "sentry-learning-to-recover-from-llm-agent-failures-at-test-time", "title": "Sentry: Learning to Recover from LLM Agent Failures at Test Time", "summary": "Sentry, an external runtime failure management layer for LLM agents released by the developer nuglifeleoji, achieved a 37% average improvement over the strongest runtime-intervention baseline across WebShop, AppWorld, SWE-bench Lite, and Mind2Web Replay, and an 81.7% recovery rate across 939 detected failures. Sentry detects invalid actions and behavioral failures, issues hard and soft repairs, verifies recovery, and turns verified soft recoveries into reusable playbook lessons, using 1.54× base-agent token usage versus up to 2.44× for other runtime methods while saving 31.7s per task on SWE-bench Lite and 45.6s on AppWorld. The tool requires Python 3.10 or later and runs alongside an existing agent loop without replacing the agent, environment, or evaluator.", "body_md": "**An external runtime failure management layer for LLM agents.**\n\nDetect execution failures, guide recovery, and learn reusable lessons from verified recoveries.\n\nSentry runs alongside your agent loop without replacing the agent, environment, or evaluator.\n\n**[Quick Start](#quick-start)**  |  **[Paper (arXiv)](https://arxiv.org/abs/2610.02994)**  |  [License](https://github.com/nuglifeleoji/Sentry/blob/main/LICENSE)\n\nSentry is an external runtime failure management layer for LLM agents. It helps agents fix invalid actions and recover when they get stuck in loops, drift from the task, or make unsupported assumptions. Unlike typical runtime monitors with fixed advice, Sentry learns and evolves from recoveries that work. By detecting failures, guiding recovery, and checking that the repair worked, Sentry helps agents become more reliable across tasks without cluttering their context.\n\n- **37% average improvement** over the strongest runtime-intervention baseline across WebShop, AppWorld, SWE-bench Lite, and Mind2Web Replay.\n- **39% average improvement** over the strongest context-evolution baseline on held-out WebShop and Mind2Web Replay tasks.\n- **1.54× base-agent token usage** , versus up to 2.44× for other runtime methods, while saving 31.7s per task on SWE-bench Lite and 45.6s on AppWorld.\n- **81.7% recovery rate** across 939 detected failures.\n\nSentry requires Python 3.10 or later.\n\n```\ngit clone https://github.com/nuglifeleoji/Sentry.git\ncd Sentry\npython -m venv .venv\nsource .venv/bin/activate\npython -m pip install -e .\nsentry-validate configs/paper.yaml\n```\n\nConfigure a model provider, then submit each completed agent-environment cycle to Sentry:\n\n``` python\nfrom Sentry import SentryRunner, agent_step_from_parts\n\nrunner = SentryRunner.from_yaml(\"configs/providers/openai_compatible.yaml\")\nrunner.start_task(\n    objective=task_objective,\n    action_schema=environment_action_schema,\n)\n\ntry:\n    step = agent_step_from_parts(\n        step_id=0,\n        reasoning=reasoning,\n        raw_action=raw_action,\n        tool_name=tool_name,\n        tool_args=tool_args,\n        observation=observation,\n        parsed_ok=True,\n        schema_valid=True,\n        accepted_by_environment=True,\n    )\n\n    repair = runner.step(step, terminal=False)\n    if repair is not None:\n        agent_messages.append(repair.prompt_text)\nfinally:\n    runner.finalize()\n```\n\nYour application remains responsible for generating, validating, and executing actions. See the integration guide for the full lifecycle and validity rules.\n\nAppWorld is the included reference integration:\n\n```\npython3.11 -m venv .venv-appworld\nsource .venv-appworld/bin/activate\npython -m pip install -e '.[appworld]'\nappworld install\nexport APPWORLD_ROOT=\"$PWD/outputs/appworld-root\"\nappworld download data\nexport OPENROUTER_API_KEY=\"your-api-key\"\nsentry run appworld --config configs/providers/openrouter.yaml\n```\n\nSee the AppWorld guide for environment setup, credentials, and run limits.\n\n1. **Failure Detection:** Sentry monitors recent agent steps for invalid actions and behavioral failures.\n2. **Hard Repair:** Sentry asks the agent to retry invalid actions in the required format.\n3. **Soft Repair:** Progress and reasoning failures receive targeted guidance from relevant playbook lessons.\n4. **Recovery Verification:** Sentry checks whether the agent recovers over the following steps.\n5. **Online Playbook Learning:** Verified soft recoveries become reusable lessons for similar failures across tasks.\n\n- **Unconditional exposure can hurt:** Controlled experiments show that keeping all failure-specific knowledge in the agent's context lowers performance. This motivates Sentry to expose recovery lessons only when a matching failure occurs.\n- **Retrieval must match the failure:** Unfiltered retrieval and broad failure-type matching both perform worse than retrieval with fine-grained failure labels, motivating Sentry's label-based retrieval.\n- **Verified lessons transfer:** Lessons learned from verified recoveries improve performance on unseen tasks, and continued learning adds further gains.\n- **Task-level and failure-level learning are complementary:** Combining Sentry with context evolution achieves the best results, showing that the two forms of learning address different needs and add on to each other.\n\nThis repository currently contains the project overview and paper. We are finalizing our code and documentation and plan to release the full version in 2-3 weeks.", "url": "https://wpnews.pro/news/sentry-learning-to-recover-from-llm-agent-failures-at-test-time", "canonical_source": "https://github.com/nuglifeleoji/Sentry", "published_at": "2026-10-06 06:26:05+00:00", "updated_at": "2026-10-06 06:49:29.021978+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools", "ai-research"], "entities": ["Sentry", "nuglifeleoji", "WebShop", "AppWorld", "SWE-bench Lite", "Mind2Web Replay", "Python 3.10", "OpenRouter"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sentry-learning-to-recover-from-llm-agent-failures-at-test-time", "markdown": "https://wpnews.pro/news/sentry-learning-to-recover-from-llm-agent-failures-at-test-time.md", "text": "https://wpnews.pro/news/sentry-learning-to-recover-from-llm-agent-failures-at-test-time.txt", "jsonld": "https://wpnews.pro/news/sentry-learning-to-recover-from-llm-agent-failures-at-test-time.jsonld"}}