Qwen UI Agent Alibaba's MAI-UI Team released Qwen-UI-Agent, a foundation GUI agent that navigates mobile, desktop, and web interfaces to complete real-world tasks, achieving 89.9 on Tau2-Bench and 64.1 on BrowseComp, outperforming Qwen3.5-27B and other models. The model preserves general reasoning and instruction-following capabilities while adding GUI-specific skills, as detailed in the Qwen-UI-Agent Technical Report. Mobile GUI Use Optimized for everyday tasks in real-device mobile GUI use, the model navigates changing Android apps, accounts, content, and network conditions to search, compare, schedule, shop, and coordinate reliably. Towards Next-Generation Real-World Centric Foundation GUI Agent One agent that thinks, searches, and acts across mobile, desktop, and the web to complete real-world, long-horizon tasks. Built to complete real work across GUI interfaces. Optimized for everyday tasks in real-device mobile GUI use, the model navigates changing Android apps, accounts, content, and network conditions to search, compare, schedule, shop, and coordinate reliably. † Author-reproduced result: the baseline was independently evaluated in the authors’ environment rather than copied from the model provider’s report. Real-world tasks demand more than interface interaction—they also require knowledge, multimodal reasoning, instruction following, and tool use. Qwen-UI-Agent gains strong GUI capabilities without becoming a narrow GUI-only model, preserving the base model’s general reasoning and agentic strengths for broader tasks. Multimodal understanding, knowledge, mathematics, and instruction following. Swipe horizontally to compare all models → | Benchmark | Qwen-UI-Agent | Qwen3.5-27B | UI-Venus 30B-A3B | GUI-Owl 32B | OpenCUA-72B | |---|---|---|---|---|---| | MMMU-Pro | 72.4 | 73.5 | 32.4 | 39.5 | 31.0 | | RealWorldQA | 83.1 | 83.1 | 75.3 | 76.7 | 66.4 | | CharXiv-RQ | 77.7 | 76.8 | 44.7 | 50.9 | 39.6 | | MathVision | 82.8 | 82.0 | 36.8 | 50.6 | 26.6 | | AI2D TEST | 91.1 | 91.9 | 84.3 | 84.8 | 78.9 | | MMLU-Pro | 86.5 | 86.0 | 65.6 | 73.9 | 58.8 | | IFEval prompt-level strict | 90.2 | 90.4 | 81.3 | 84.5 | 70.6 | Tool use, terminal tasks, multi-turn service work, coding workflows, and deep-research retrieval on BrowseComp BC and BrowseComp-ZH BC-ZH . Swipe horizontally to compare all models → | Benchmark | Qwen-UI-Agent | Qwen3.5-27B | UI-Venus 30B-A3B | GUI-Owl 32B | OpenCUA-72B | |---|---|---|---|---|---| | Tau2-Bench | 89.9 | 89.2 | 22.7 | 6.1 | 14.4 | | Terminal-Bench 2.0 · Avg 5 | 50.1 | 41.1 | 3.2 | 0.0 | 9.0 | | Claw-Eval · Avg 3 | 73.5 | 66.9 | 30.6 | 29.6 | 26.4 | | Claw-Eval · Pass@3 | 51.8 | 41.2 | 5.5 | 5.5 | 0.5 | | BFCL-v4 | 74.2 | 71.3 | 19.8 | 32.7 | 28.3 | | SkillsBench · Avg 5 | 28.0 | 24.9 | 0.5 | 0.3 | 0.0 | | QwenClawBench · Avg 3 | 44.2 | 48.5 | 6.4 | 5.1 | 11.4 | | BrowseComp BC | 64.1 | 61.0 | — | — | — | | BrowseComp-ZH BC-ZH | 75.0 | 62.1 | — | — | — | All scores were independently reproduced in the authors' evaluation environment. Some harness, judge, simulator, runtime, or task-subset settings differ from official evaluations and are documented in the technical report. Recipe research + e-shopping I’m planning to make “passion-fruit sour-soup beef” tonight. Search Douyin for the most-saved photo-and-text post, save it, and remember the ingredients I need to prepare. Then, in the Hema app, purchase all the ingredients mentioned in the post—excluding seasonings—select delivery for 18:45 today, and place the order. Complete everyday workflows across changing apps on physical mobile devices. The agent first extracts a recipe and its ingredients from Douyin, then carries that information into Hema to complete a time-constrained grocery order. @misc{qwenuiagent2026, title = {Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World-Centric Foundation GUI Agents}, author = {MAI-UI Team}, year = {2026}, note = {Alibaba Token Hub} }