GitHub Security Lab says its open-source workflows found and reported 24 flaws. The examples include an OsmAnd location-tracking bug and a Wikipedia app deep-link issue; running the mobile audit requires a Copilot license and premium model requests.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [The GitHub Blog](https://github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-security-agent/)
Why it matters #
A reusable audit workflow can help security teams examine mobile-specific attack surfaces that general code review may miss, including exported Android components and deep links. GitHub's examples also show the operational tradeoff: the workflows are inspectable and runnable, but depend on premium model requests and can consume significant time and tokens.
GitHub Security Lab says it found and reported 24 vulnerabilities in Android applications using custom AI audit workflows. In a post dated September 28th and published September 29th, researcher Kevin Stubbings described two disclosed examples and published the taskflows for others to run. The report is GitHub Security Lab's account of a system it built, rather than an independent evaluation of its findings. GitHub's account
The first example involves OsmAnd, a navigation app with more than 10 million Play Store downloads, according to GitHub. Its exported MapActivity accepts intent extras used when importing settings. Because another app can send arbitrary extras to an exported activity, GitHub says a malicious app could trigger an import without a notification or user confirmation, then replace settings that control map tiles. The attacker could route tile requests through a server they control and infer the coordinates of tiles the user viewed. GitHub also describes capturing route origins and destinations through the same issue. Its post says OsmAnd had three vulnerabilities; the location-tracking flaw is the one discussed in detail.
The second example concerns the Wikipedia Android app's wikipedia:// deep links. GitHub says a hostname-parsing bug let a link open a non-Wikipedia URL inside the app, where an attacker-controlled page could run JavaScript in its WebView. The available text of the post cuts off while introducing a second code snippet about cookie handling. It does not provide enough information to verify the subsequent cookie and token-exposure chain or the stated testing limitations, so those details are not described here.
The taskflows are YAML-defined sequences of tasks run by GitHub Security Lab's agent framework. The framework gives each task a prompt and designated tools, and can pass results between tasks through a toolbox such as a memory cache. GitHub says each task starts with a fresh context, which makes task-by-task outputs easier to rerun while debugging. The framework itself is described in an earlier GitHub Security Lab post; the companion seclab-taskflows repository contains the example workflows and supporting MCP servers.
For the Android audit, Stubbings added gather_mobile_entry_point_info.yaml to separate mobile entry points from other entry points in a repository. He also modified classify_application_local.yaml to prompt the model to check for specified vulnerability classes in each entry point and component. For example, when a component uses an Android intent, the taskflow checks for issues such as a confused deputy or insecure broadcasts. GitHub says it runs both narrower, strict prompts and broader prompts multiple times, aiming to catch expected classes of bugs while leaving room for the model to identify other issues.
To try the mobile audit, start a Codespace from the taskflows repository, allow it to initialize, then run ./scripts/audit/run_mobile.sh myorg/myrepo in the terminal. GitHub says a medium-sized repository can take an hour or two and that results open in an SQLite viewer; its post directs users to the audit_results table and rows marked in the has_vulnerability column. A Copilot license is required, and the prompts use premium model requests. The run can make many tool calls and consume substantial tokens.
The repository also documents local and container-backed execution requirements: Python 3.11 or later for local runs, Docker for taskflows that use container-backed tools, and configured AI API credentials and endpoint variables. Its README warns that audits can take several hours, especially on larger projects, and make enough AI requests to incur a non-trivial cost. GitHub's framework post describes the project as experimental and separates the agent implementation from the taskflow suite, so the workflows can be inspected and adapted rather than treated as a closed security scanner.