{"slug": "the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm", "title": "The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents", "summary": "A new arXiv survey (2608.06663v1) of 1,547 papers (2024-2026) identifies a 'horizon gap' where frontier language models fail at multi-hour tasks despite solving single-pass reasoning problems, and proposes a taxonomy distinguishing long-horizon, long-context, and long-term memory. The authors organize the corpus into six lifecycle categories and argue that outcome-only signals become uninformative as horizons lengthen, highlighting open measurement problems in model versus harness capability and correlated bias in process-level signals.", "body_md": "arXiv:2608.06663v1 Announce Type: new\nAbstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.", "url": "https://wpnews.pro/news/the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm", "canonical_source": "https://arxiv.org/abs/2608.06663", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 04:11:42.039010+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm", "markdown": "https://wpnews.pro/news/the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm.md", "text": "https://wpnews.pro/news/the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm.txt", "jsonld": "https://wpnews.pro/news/the-horizon-gap-planning-memory-execution-training-and-evaluation-for-long-llm.jsonld"}}