cd /news/ai-agents/artemis-google-s-new-ai-agent-framew… · home topics ai-agents article
[ARTICLE · art-127491] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Artemis: Google's new AI agent framework for mobile test automation

Google released Artemis, an open-source AI agent framework that automates Android testing by driving real phones from natural-language instructions, achieving 99%+ task completion on Google Research's AndroidWorld benchmark of 100+ multi-step tasks. Artemis integrates a native Model Context Protocol (MCP) server so AI IDEs including Antigravity, Claude Code, Windsurf, Cursor, Codex, VS Code, Cline/Roo, and OpenClaw can control test devices and collect Logcat output and screenshots, with a reactive observe-and-act loop running at typically 3–5 seconds per step. The framework installs ADB, scrcpy, FFmpeg, and Python (uv) dependencies via a one-click start.sh or start.bat script and serves a local interface at http://localhost:8000.

read9 min views1 publishedSep 12, 2026
Artemis: Google's new AI agent framework for mobile test automation
Image: Michielbdejong (auto-discovered)

Let AI assistants and test suites use real phones like a human.

English中文文档Workflow ShowcaseQuick StartMCP for IDEsBenchmarksDiscord Community

Live Demo: Setup driving routes and calculate total durations in Google Maps, then open YouTube to play a Coldplay song.

  • Cross-App Automation : Executes testing workflows and everyday tasks on Android from natural language instructions.
  • Multimodal Targeting : Uses element indices when available, with coordinate and visual locating fallbacks for custom interfaces.
  • IDE Diagnostics :Model Context Protocol (MCP) integration letsAntigravity, Claude Code, and Windsurf drive test devices and collectLogcat output and screenshots.
  • Flash Execution : A reactive observe-and-act loop with asynchronous history summaries, typically3–5s per step .
  • Pro Exploration : Checks targets before individual actions and returns blocked actions to the Operator for recovery. Supports long-running exploratory and stability tests.
  • AndroidWorld Results :99%+ task completion on Google Research'sAndroidWorld benchmark (100+ multi-step tasks).

Antigravity uses ARTEMIS through MCP to turn a test request into a plan, device execution, and a diagnostic report:

| 1. Prompt Input (Task Dispatch)<sub>Describe your test scenario and target metrics in Antigravity</sub> | 2. Test Plan Generation<sub>Formulates a step-by-step test plan & architecture for review</sub> | | 3. Autonomous Test Execution<sub>Drives real device, navigates UI, and profiles performance</sub> | 4. Final Report<sub>Delivers structured audit findings, metric tables, and raw datasets</sub> |

Ensure an Android device (with USB Debugging enabled) or emulator is connected. The one-click startup script will automatically:

  • Install System Toolchains : Detect and auto-install ADB, scrcpy, FFmpeg, and Python (uv ) dependencies.
  • Mount Global MCP Server & AI Agent Rules : Prompt to automatically install global MCP configurations and theArtemis Mobile Testing Mindset (rules.md) into your AI IDEs (Antigravity ,Cursor ,Claude Code ,Codex ,Windsurf ,VS Code ,Cline/Roo ,OpenClaw ).
git clone https://github.com/google/artemis.git && cd artemis

./start.sh
git clone https://github.com/google/artemis.git
cd artemis

.\start.bat

PowerShell does not search the current directory for executable scripts by default, so use .\start.bat without a trailing \. In Command Prompt (CMD), use start.bat instead.

Tip: Opens http://localhost:8000 in your default browser with a device connection wizard, live screen mirroring, prompt sandbox, and execution replays. You can also run directly from CLI: uv run artemis run "Open Settings, find Battery and tell me current level" --profile flash.

MCP Setup for Codex / Antigravity / Claude Code / Windsurf (Click to expand) #

ARTEMIS includes a native Model Context Protocol (MCP) server. Connect your real phone directly into AI IDEs:

Running ./start.sh (macOS/Linux) or .\start.bat (Windows PowerShell) will prompt you to configure global MCP and testing rules for detected IDEs (or you can install/update anytime later manually using the commands below):

uv run artemis mcp --install antigravity

uv run artemis mcp --install all

Tip: You can also configure MCP interactively during first-time setup via uv run artemis init. Pro Tip: If you want to use the artemis command globally without uv run in any directory, run uv tool install -e . once in the project root.

If you prefer to configure manually, run uv run artemis mcp --generate-config <client> (for example, codex or antigravity) to output the appropriate TOML or JSON snippet. Replace /path/to/artemis with your actual repo path and point command to your .venv Python executable:

  • Codex (~/.codex/config.toml ):
[mcp_servers.artemis]
command = "/path/to/artemis/.venv/bin/python"
args = ["-m", "mcp_server"]
cwd = "/path/to/artemis"

[mcp_servers.artemis.env]
PYTHONUNBUFFERED = "1"
PYTHONPATH = "/path/to/artemis"
  • Antigravity (~/.gemini/jetski/mcp_config.json ):
{
  "mcpServers": {
    "artemis": {
      "command": "/path/to/artemis/.venv/bin/python",
      "args": ["-m", "mcp_server"],
      "cwd": "/path/to/artemis",
      "env": {
        "PYTHONUNBUFFERED": "1"
      },
      "tools": {
        "mobile_run_task": { "eager": true },
        "mobile_manage_task": { "eager": true },
        "mobile_get_device_state": { "eager": true },
        "mobile_inspect_trace": { "eager": true },
        "mobile_diagnose": { "eager": true }
      }
    }
  }
}
  • Claude Desktop (claude_desktop_config.json ):
{
  "mcpServers": {
    "artemis": {
      "command": "/path/to/artemis/.venv/bin/python",
      "args": ["-m", "mcp_server"],
      "cwd": "/path/to/artemis"
    }
  }
}

To ensure your AI coding assistant acts with the rigor of a senior mobile test engineer and never hallucinates UI interactions, we provide a dedicated testing mindset rules file at mcp_server/rules.md (covering Active Exploration before coding, Flash vs. Pro routing strategy, Latency & Timing compensation, and the "Dynamic-First, Coordinate-Fallback" locator pattern).

You can mount or copy mcp_server/rules.md into your AI IDE's rule configuration:

  • Antigravity : Add the contents ofrules.md to your Workspace Rules, Global Rules settings, or agent instructions.
  • Claude Code : Runartemis mcp --install claude to install the rules to~/.claude/rules/artemis.md (install to exactly one location — Claude Code loads both~/.claude/CLAUDE.md and~/.claude/rules/*.md , so duplicating the rules wastes context).
  • Cursor : Copy the contents into.cursorrules or create a rule file at.cursor/rules/artemis.mdc .
  • Codex : Add the contents to~/.codex/AGENTS.md (or the activeAGENTS.override.md ).
  • Windsurf / OpenClaw : Add the rules to your workspace rules or global system prompts.

For more details on the testing mindset and MCP architecture, see the MCP Server README.

In Codex, Antigravity, or Claude Code, simply prompt:

"Build the latest changes into an APK, install it on the connected device, open the login screen with a test account, verify if there are any unexpected popups after login, and return screenshots of the final page."

Python SDK Integration (Click to expand) #

Install the zero-runtime-dependency client on the development machine. ADB, agents, models, and image processing remain on the device host:

uv add "artemis-client @ git+https://github.com/google/artemis.git#subdirectory=packages/artemis-client"
python
import asyncio
from artemis_client import ArtemisClient

async def main():
    client = ArtemisClient(
        "http://artemis-host:8000",
        device_serial="emulator-5554",  # optional: target specific device serial
        default_profile="flash",  # "flash" (fast reactive) or "pro" (deep reasoning)
    )

    result = await client.run(
        "Open System Settings, go to 'Battery', verify battery percentage is displayed, and check for any crash dialogs.",
    )

    assert result.succeeded, f"Test failed: {result.error or result.status}"
    print(f"✅ Test Passed! Device: {result.device_serial} | Trace ID: {result.trace_id}")

if __name__ == "__main__":
    asyncio.run(main())

Console Overview: ① View Switcher (Home / Workspace) · ② Model & Replay (Flash/Pro status & video replay) · ③ Live Agent Stream (Action perception, target coordinates & structured results) · ④ Prompt Dock (Natural language dispatch) · ⑤ Task Queue & Dashboard (Lifecycle & history)

  • Web Visual Test Console (uv run artemis ui) : Real-time screen projection and interactive panel, supporting natural language test dispatch, live reasoning telemetry, action trajectories, and execution replay; manage server lifecycle anytime from any terminal usinguv run artemis restart ,uv run artemis stop , anduv run artemis status ;
  • MCP Server : ConnectsAntigravity, Claude Code, Windsurf , and other MCP clients to real devices for bug reproduction and test execution;
  • Developer CLI (uv run artemis run) : Direct terminal execution for automated test cases, exploratory stability inspection, or AndroidWorld benchmarks with high-fidelity structured terminal output;
  • Python SDK : Integrates as a standard Python library into existing automated testing frameworks (e.g., pytest) or CI/CD pipelines with strongly typed Pydantic structured outputs and assertion support.

The first task on a device installs the Artemis Accessibility Helper, a small accessibility service that reads the screen layout without taking the UiAutomation connection. Tools using UiAutomation can suppress the helper unless they enable FLAG_DONT_SUPPRESS_ACCESSIBILITY_SERVICES. You will see a collapsed "Artemis test helper is running" notification and a new entry under Settings > Accessibility; both are that helper. It listens only on the phone itself and sends nothing elsewhere.

  • Pre-install it (avoids the ~3 s delay on the first task): uv run artemis helper install
  • Inspect it: uv run artemis helper status /uv run artemis doctor
  • Remove it any time: uv run artemis helper uninstall
  • Use UIAutomator2 instead: ARTEMIS_HIERARCHY_BACKEND=uiautomator in.env
  • Prevent automatic installation: ARTEMIS_HELPER_AUTO_INSTALL=false in.env

If the helper ever fails mid-task, ARTEMIS falls back to UIAutomator2 and says so in the task timeline, in mobile_manage_task status, and in the final report.

Artemis achieved a 99%+ completion rate on AndroidWorld, Google Research's benchmark spanning 20+ apps and 100+ multi-step tasks.

  • Pre-Execution Checks and Action Bursts : Pro checks the target against the live UI tree and pixels before dispatching an individual action. Action bursts handle transient controls without waiting for another model turn.
  • Element Locating : Combines accessibility hierarchies and OCR with visual models for custom Canvas, Compose, and Flutter interfaces.
  • Shared History Compression : Flash and Pro replace older screenshots with visual summaries and compress completed steps into searchable history chunks. Context thresholds control when raw turns are replaced.

ARTEMIS supports two execution profiles tailored for different automation requirements:

  • Flash Profile (--profile flash) : Fast and token-efficient reactive loop (~3–5s per step): one model observes the live screen, thinks, and acts, with no graph orchestration. Ideal for routine, deterministic UI tasks. The loop is unbounded by default (agent.flash.max_turns , 0 = unlimited) because history is compressed rather than capped: Flash shares the Pro session transcript ledger (session-relativeT+mm:ss clock, screenshots folded into visual summaries, older steps chunked into eras and recallable on demand viasearch_history /replay_steps ) and can query the session recording throughvideo_analyzer . Transient UI (auto-fading control bars, toasts) is handled by chaining taps into oneclick_sequence .Limitations : No task plan or notes, no pre-execution safety net, no checkpoint verification or final report, and no ADB shell.

  • Pro Profile (--profile pro) : A planning and verification workflow (~15–40s per step), built as a multi-agent graph. APlanner maintains a living Markdown task plan with milestones andverify /assert check items; theOperator executes it with the full toolset (Explorer grounding whoseflash /pro /ultra tier is a user setting per profile —pro.explorer.mode /flash.explorer_mode inconfig/artemis.jsonc or--explorer-pro-mode — never chosen by the agent; notes, history recall, video analysis, ADB diagnostics). Every single action passes a pre-executionSafety Net (XML-first, pixel fallback), while multi-actionfast-action bursts fire back to back to beat turn latency on transient UI. A blocked or failed action opens anexecution incident that stays in the Operator's context until a later action succeeds, so recovery is handled by the Operator itself with no separate repair agent. A read-onlyChecker verifies plan checkpoints and runs an exit final review against the original goal (--verification-level :off /final (default) /checkpoints /strict ), and plan milestone edits get an advisory review. Handles 100+ step long-horizon workflows,[Loop:continuous] monitoring, and an optional written report.

  • Android Studio Integration : Native IDE plugin and workflow integration to enable in-editor debugging, test recording, and automated device control directly within Android Studio.

  • iOS Platform Expansion : Extending multimodal perception and mobile automation to iOS devices and simulators.

  • On-Device Lightweight VLMs : Local execution with lightweight edge vision models for low-latency, privacy-first automation.

  • Real-time Duplex Voice Interaction : Voice-driven task dispatch with real-time conversational control and interruption handling.

Contributions are warmly welcomed!

This project is licensed under the Apache License 2.0.

This project includes source code developed by Minitap, Inc..

── more in #ai-agents 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/artemis-google-s-new…] indexed:0 read:9min 2026-09-12 ·