{"slug": "learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural", "title": "Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers", "summary": "Researchers introduced rollout-verified policy improvement (RV-PI), a long-horizon reinforcement learning method for budgeted neural-operator PDE solvers, detailed in arXiv paper 2610.06883v1. On the shallow-water benchmark with a 32-intervention budget, RV-PI reached a three-seed mean trajectory relative L2 error of 0.6910, improving 5.37% over immediate-only policy improvement and 2.41% over RandomMacro; on the forcing-driven Brusselator benchmark with a 76-intervention budget it attained 0.09954, improving 2.31% and 5.32% respectively. The results indicate that under a fixed refinement budget, the value of a local correction depends on its downstream effect on the autoregressive trajectory, not only its immediate error reduction.", "body_md": "arXiv:2610.06883v1 Announce Type: new \nAbstract: Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections can be committed along a trajectory. We formulate this as budgeted adaptive neural-operator solving. A global Fourier neural operator advances the full field, a local operator proposes patch-wise residual corrections, and a set-aware selector chooses where to refine. A macro policy decides when and how much of the remaining refinement budget to spend. We introduce rollout-verified policy improvement (RV-PI), which evaluates feasible refinement counts through actual continuation rollouts of the learned PDE solver, converts long-horizon advantages into conservative policy targets, and accepts an update only when held-out trajectory error improves. On the shallow-water benchmark with a 32-intervention budget, RV-PI achieves a three-seed mean trajectory relative L2 error of 0.6910, improving over immediate-only policy improvement by 5.37% and RandomMacro by 2.41%. On the forcing-driven Brusselator benchmark with a 76-intervention budget, RV-PI attains 0.09954, improving over immediate-only policy improvement by 2.31% and RandomMacro by 5.32%. These results show that, under a fixed refinement budget, the value of a local correction depends on its downstream effect on the autoregressive trajectory, not only on its immediate error reduction.", "url": "https://wpnews.pro/news/learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural", "canonical_source": "https://arxiv.org/abs/2610.06883", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:16:42.631533+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "neural-networks"], "entities": ["RV-PI", "Fourier neural operator", "shallow-water benchmark", "Brusselator benchmark", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural", "markdown": "https://wpnews.pro/news/learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural.md", "text": "https://wpnews.pro/news/learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural.txt", "jsonld": "https://wpnews.pro/news/learning-when-to-refine-long-horizon-reinforcement-learning-for-budgeted-neural.jsonld"}}