The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents A new arXiv survey (2608.06663v1) of 1,547 papers (2024-2026) identifies a 'horizon gap' where frontier language models fail at multi-hour tasks despite solving single-pass reasoning problems, and proposes a taxonomy distinguishing long-horizon, long-context, and long-term memory. The authors organize the corpus into six lifecycle categories and argue that outcome-only signals become uninformative as horizons lengthen, highlighting open measurement problems in model versus harness capability and correlated bias in process-level signals. arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers 2024-2026 collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon task property: required steps , long-context model property: token capacity , and long-term memory system property: persistence across steps/sessions . We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried within-context, within-task-beyond-context, or cross-task-persistent . Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.