cd /news/artificial-intelligence/learn-now-use-next-trust-later-prequ… · home › topics › artificial-intelligence › article
[ARTICLE · art-142236] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

A new arXiv paper introduces StepLearn, a nonparametric framework that lets large language model agents learn from individual transitions during deployment while requiring prospective validation before rules are reused across episodes. Over five rounds on WebArena-Lite and ALFWorld, StepLearn reached average success rates of 59.9% and 84.0% with GPT-5-mini and 57.8% and 88.1% with Qwen3.5-35B-A3B, beating the strongest baseline, EvoTest, by 2.2 to 12.7 percentage points across the four settings. The paper reports that gains appeared on first task attempts in most settings, not only on final repetitions, with all model parameters kept fixed.

by read1 min views2 publishedSep 30, 2026

arXiv:2609.35911v1 Announce Type: new Abstract: Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @steplearn 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/learn-now-use-next-t…] indexed:0 read:1min 2026-09-30 · —