GitHub Security Lab researcher Kevin Stubbings wrote on September 28 about how his team used GitHub's own open-source AI security agent to audit Android apps, finding and reporting 24 vulnerabilities in the process. The tool, called seclab-taskflow-agent, is built around a set of YAML-defined "task flows" that walk a large language model step by step through reading code, classifying entry points, spotting issues, and writing proof-of-concept code.
The task flows had previously been used for general code auditing; this time the team customized them for mobile, adding two new files — gather_mobile_entry_point_info.yaml and classify_application_local.yaml — to handle mapping an app's entry points and categorizing vulnerabilities.
Getting it running
Per the writeup, users run an audit script inside GitHub Codespaces, wait a few minutes for the environment to initialize, then let the agent work. A mid-sized repository takes roughly one to two hours. Results land in a SQLite database; opening the audit_results table and filtering for rows where has_vulnerability is checked shows what the agent flagged as a problem.
The task flows target a familiar set of Android weak points: Intent-based confused-deputy issues and insecure broadcasts, cross-app scripting in WebViews, JavaScript bridges exposed to web content, path traversal, flawed deep-link parsing, and leaked cookies or login tokens. The writeup says the classification list spans 12 CWE categories.
Two cases in detail
OsmAnd, an open-source map and navigation app with more than 10 million downloads, had three issues in its settings-import flow. One of them lets a malicious app on the same phone silently import configuration via Intent extras — without requesting any permissions — and use it to track the user's location and routes.
The Wikipedia Android app had a flaw in its deep-link parsing logic. An attacker could craft a malicious link that leads the user to a spoofed page, use it to steal cookies, and ultimately take over the account with a long-lived login token.
Stubbings wrote that the team was surprised by how accurately the model understood security-relevant API behavior across different languages, even without source code in that language; the proof-of-concept code the agent produced needed only minor changes to actually run.
"LLMs are good at finding vulnerabilities but struggle at estimating severity."
Where it falls short
The writeup is candid about the limits. The model often flags low-severity issues even when the prompt explicitly asks it not to; when a vulnerability has mitigations in place, it struggles to judge the real-world impact; it misses complex data-flow prioritization problems; and getting a proof of concept that actually works often takes several runs, sometimes with a debugger attached or extra prompting. The conclusion: every finding still needs manual review by a researcher who knows mobile security.
On cost, running the agent requires a GitHub Copilot license and uses premium model requests. The writeup warns that large repositories generate a lot of tool calls and burn through tokens quickly; it doesn't specify which model was used.
For developers in China
The task flows and scripts are open source on GitHub, so teams in China could adapt them to audit their own Android apps — the main obstacles being the Copilot license and reliance on overseas models. The task flows themselves are just YAML-defined prompts and step sequencing, so swapping in a domestic model is theoretically workable, though there's no published comparison of how well that would perform.
Android's fragmentation makes this kind of tooling especially useful in China. The WebViews and JavaScript bridges embedded in phone makers' app stores, mini-program containers, and super-apps fall squarely within the vulnerability categories this checklist covers. GitHub says it plans to extend the task flows to web and desktop apps next, and open the project up to community contributions.
Sources: GitHub's official blog, CocoLoop, GitHub Security Lab's open-source repository; vulnerability counts, OsmAnd's download figure, and per-audit runtime are as stated in GitHub's blog post.