{"slug": "android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations", "title": "Android Bench 2.0 focuses on long-horizon tasks, agent evaluations", "summary": "Google released Android Bench 2.0, a benchmark that shifts from binary pass/fail grading to continuous completion-rate scoring for long-horizon Android development tasks that take an engineer multiple days or a week. OpenAI's GPT-6 Astra tops the new benchmark with a 28% pass rate, compared with scores in the 90% range under the previous approach, and Google reported that no model reaches a 100% pass rate on porting cross-platform apps to Android while frontier models reach at most an 80% completion rate. Google rated Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max, and ran agents from the corresponding model providers, including Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex.", "body_md": "Google’s development of Android Bench [continues today](https://android-developers.googleblog.com/2026/09/android-bench-2-long-horizon-tasks.html) with a version 2.0 that reflects how AI can handle more complex development tasks.\n\nThe first version “focused on incremental changes to existing repositories,” like bug fixes or smaller feature requests. [Android Bench 2.0](https://developer.android.com/bench) targets “tasks of great complexity that take an engineer multiple days or even a week to complete,” such as adding new features, building apps from scratch, and converting cross-platform apps to Android.\n\nThis new focus required “more nuanced evaluation and scoring” that goes beyond pass or fail grading. Google is moving from binary to continuous scoring:\n\nWe calculate this completion rate through a combination of factors like functionality, visual fidelity, and avoiding regressions. We also apply objective scoring penalties for deviations from evaluation instructions or structural constraints.\n\nGoogle has rated Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. GPT-6 Astra is at the top of the benchmark with a 28% pass rate (compared to scores in the 90% range with the previous approach).\n\nGoogle shared insights like how “porting cross-platform apps to Android remains an open challenge—no model hits a 100% pass rate, and frontier models reach at most a 80% completion rate.”\n\n- “…AI does a better job at writing new code rather than refactoring existing code. Refactors and migrations get trickier because success depends on architectural complexity rather than code volume.”\n- “Models show strong capabilities on well-established, deterministic transformations, such as converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer.”\n- “…models struggle when tasks require runtime validation (like missing dependency injection graphs), involve breaking framework changes, or run into knowledge gaps with unreleased libraries.”\n\nWith agent evaluations, Android Bench ran “agents from the corresponding model provider,” like Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex. Google says “harness design positively impacts developer outcomes,” with Android Bench planning to include different model and agent combinations in the future.\n\n*FTC: We use income earning auto affiliate links.* [More.](https://9to5mac.com/about/#affiliate)\n\n[our homepage](http://9to5google.com/)for all the latest news, and follow 9to5Google on\n\n[exclusive stories](https://9to5google.com/feature/exclusive/),\n\n[reviews](https://9to5google.com/guides/reviews/),\n\n[how-tos](https://9to5google.com/guides/how-to/), and\n\n[subscribe to our YouTube channel](https://www.youtube.com/9to5google)", "url": "https://wpnews.pro/news/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations", "canonical_source": "https://9to5google.com/2026/09/17/android-bench-2-0/", "published_at": "2026-09-17 16:00:00+00:00", "updated_at": "2026-09-17 16:26:55.400675+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "large-language-models", "developer-tools"], "entities": ["Google", "Android Bench", "Gemini 3.7/3.8 Flash", "OpenAI", "GPT-6 Astra", "Anthropic", "Fable 5.1", "Google Antigravity"], "alternates": {"html": "https://wpnews.pro/news/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations", "markdown": "https://wpnews.pro/news/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations.md", "text": "https://wpnews.pro/news/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations.txt", "jsonld": "https://wpnews.pro/news/android-bench-2-0-focuses-on-long-horizon-tasks-agent-evaluations.jsonld"}}