cd /news/artificial-intelligence/android-bench-2-0-focuses-on-long-ho… · home topics artificial-intelligence article
[ARTICLE · art-132782] src=9to5google.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Android Bench 2.0 focuses on long-horizon tasks, agent evaluations

Google released Android Bench 2.0, a benchmark that shifts from binary pass/fail grading to continuous completion-rate scoring for long-horizon Android development tasks that take an engineer multiple days or a week. OpenAI's GPT-6 Astra tops the new benchmark with a 28% pass rate, compared with scores in the 90% range under the previous approach, and Google reported that no model reaches a 100% pass rate on porting cross-platform apps to Android while frontier models reach at most an 80% completion rate. Google rated Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max, and ran agents from the corresponding model providers, including Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex.

by read2 min views3 publishedSep 17, 2026
Android Bench 2.0 focuses on long-horizon tasks, agent evaluations
Image: 9To5Google (auto-discovered)

Google’s development of Android Bench continues today with a version 2.0 that reflects how AI can handle more complex development tasks.

The first version “focused on incremental changes to existing repositories,” like bug fixes or smaller feature requests. Android Bench 2.0 targets “tasks of great complexity that take an engineer multiple days or even a week to complete,” such as adding new features, building apps from scratch, and converting cross-platform apps to Android.

This new focus required “more nuanced evaluation and scoring” that goes beyond pass or fail grading. Google is moving from binary to continuous scoring:

We calculate this completion rate through a combination of factors like functionality, visual fidelity, and avoiding regressions. We also apply objective scoring penalties for deviations from evaluation instructions or structural constraints.

Google has rated Gemini 3.7/3.8 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. GPT-6 Astra is at the top of the benchmark with a 28% pass rate (compared to scores in the 90% range with the previous approach).

Google shared insights like how “porting cross-platform apps to Android remains an open challenge—no model hits a 100% pass rate, and frontier models reach at most a 80% completion rate.”

  • “…AI does a better job at writing new code rather than refactoring existing code. Refactors and migrations get trickier because success depends on architectural complexity rather than code volume.”
  • “Models show strong capabilities on well-established, deterministic transformations, such as converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer.”
  • “…models struggle when tasks require runtime validation (like missing dependency injection graphs), involve breaking framework changes, or run into knowledge gaps with unreleased libraries.”

With agent evaluations, Android Bench ran “agents from the corresponding model provider,” like Gemini 3.8 Flash on Google Antigravity and GPT-5.6 Sol with Codex. Google says “harness design positively impacts developer outcomes,” with Android Bench planning to include different model and agent combinations in the future.

*FTC: We use income earning auto affiliate links.* [More.](https://9to5mac.com/about/#affiliate)

[our homepage](http://9to5google.com/)for all the latest news, and follow 9to5Google on

[exclusive stories](https://9to5google.com/feature/exclusive/),

[reviews](https://9to5google.com/guides/reviews/),

[how-tos](https://9to5google.com/guides/how-to/), and

[subscribe to our YouTube channel](https://www.youtube.com/9to5google)
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/android-bench-2-0-fo…] indexed:0 read:2min 2026-09-17 ·