{"slug": "beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon", "title": "Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents", "summary": "Researchers identified a failure mode they call Decision-Timestamp Mismatch in on-policy self-distillation (OPSD), a technique used to add dense privileged feedback to reinforcement learning with verifiable rewards (RLVR) for long-horizon agents. The finding addresses how RLVR's sparse outcome rewards provide only coarse supervision, while OPSD's dense feedback can misalign decisions with their timestamps.", "body_md": "Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged", "url": "https://wpnews.pro/news/beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon", "canonical_source": "https://aiflash.com/news/128222/", "published_at": "2026-09-29 03:00:01+00:00", "updated_at": "2026-09-29 03:18:50.353897+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-agents"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon", "markdown": "https://wpnews.pro/news/beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon.md", "text": "https://wpnews.pro/news/beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon.txt", "jsonld": "https://wpnews.pro/news/beyond-timestamps-decision-aligned-on-policy-distillation-for-long-horizon.jsonld"}}