{"slug": "beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents", "title": "Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents", "summary": "Researchers introduced LACL-GUI, a length-aware contrastive learning framework for GUI agents powered by multimodal large language models, which improves reinforcement learning stability and performance by incorporating trajectory-level quality signals. In experiments on GUI agent benchmarks, LACL-GUI consistently outperformed prior methods, demonstrating the value of fine-grained supervision in contrastive RLVR.", "body_md": "arXiv:2608.21830v1 Announce Type: new\nAbstract: Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks across diverse digital environments, where reinforcement learning (RL) has become a dominant training paradigm. However, widely used methods such as Group Relative Policy Optimization (GRPO) suffer from reward-gradient misalignment, leading to inefficient and unstable optimization. Recent work addresses this issue by reformulating RL with verifiable rewards (RLVR) as contrastive or classification-based objectives, which improve stability by eliminating problematic gradient behaviors. Despite this progress, existing contrastive RLVR methods rely primarily on outcome-level supervision and fail to capture fine-grained differences in trajectory quality within the same outcome category. In this paper, we propose Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive RLVR framework that incorporates trajectory-level quality signals into policy optimization. LACL-GUI introduces structured preferences within both successful and failed trajectories, encouraging concise successful executions and differentiating failure quality based on divergence from successful trajectories, while preserving optimization stability. Experiments on GUI agent benchmarks show that LACL-GUI provides more effective learning signals and consistently improves agent performance over prior methods, highlighting the value of trajectory-level supervision in contrastive RLVR.", "url": "https://wpnews.pro/news/beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents", "canonical_source": "https://www.machinebrief.com/news/beyond-success-and-failure-length-aware-contrastive-learning-1ybk", "published_at": "2026-08-25 04:00:00+00:00", "updated_at": "2026-08-25 05:13:58.135754+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-agents", "ai-research"], "entities": ["LACL-GUI", "arXiv", "Multimodal Large Language Models", "Group Relative Policy Optimization", "Reinforcement Learning with Verifiable Rewards"], "alternates": {"html": "https://wpnews.pro/news/beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents", "markdown": "https://wpnews.pro/news/beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents.md", "text": "https://wpnews.pro/news/beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents.txt", "jsonld": "https://wpnews.pro/news/beyond-success-and-failure-length-aware-contrastive-learning-for-gui-agents.jsonld"}}