Uncovering AI Vulnerabilities: Insights from Google's Agent Development Kit Security Flaws
Recent security vulnerabilities discovered in the GitHub repository for Google's Agent Development Kit (ADK) for Python have illuminated significant risks associated with automated AI workflows. According to a report from Pillar Security, these flaws could enable public-facing AI agents to execute unauthorized actions, potentially manipulating pull-request reviews and exposing sensitive credentials.
The first exploit path involves a triage agent that assesses pull requests submitted by external contributors. This agent, using the adk-bot account—which has collaborator permissions—can be manipulated via malicious commands hidden in pull requests. These commands can force the agent to invoke an “@gemini-cli” command, activating workflows restricted to trusted users.
While this workflow does not allow code pushing via the associated GitHub token, it does permit alterations to issues and pull requests. Essentially, through strategic exploitation, an attacker could change a maintainer's comment, submit a fake approval as github-actions[bot], or even eliminate legitimate review requests, making malicious pull requests appear merge-ready.
Pillar Security successfully recreated this attack vector in their research setup. However, they noted that a maintainer's final approval is still required to execute a merge. Subsequent to these findings, Google took measures to enhance the security of the affected repository.
A different vulnerability emerged linked to a newer Antigravity-based agent. Here, attackers could inject prompts via public issues, leading an analysis agent to execute commands meant for reliable users. Although the fixing workflow aimed to restrict commands to Git and GitHub only, the vulnerabilities identified allowed for arbitrary code execution. Researchers were able to extract the adk-bot's personal access token, as well as a Google Cloud service account key, which were both accessible to the workflow—and exploitable by attackers.
Pillar Security confirmed that as of July 2, these compromised workflows were removed from the repository, and Google informed them on July 21 that the second flaw was also patched.
Agent Handoffs Expose Risks
Pillar referred to its findings as a noteworthy instance of agent-to-agent exploitation in operational multi-agent systems. According to Sanchit Vir Gogia, chief analyst at Greyhound Research, while the vulnerabilities are not new, how they interact requires companies to reassess the authority dynamics within agent systems.
“Natural language has entered the authorization pathway,” Gogia commented, emphasizing that what deserves attention is not merely the novelty of this situation but the implications it carries for security protocols.
Gogia's perspective suggests that authority granted to an agent shouldn’t only be evaluated based on its toolset but also on how its outputs can influence more privileged systems. This broader understanding could dictate the way Chief Information Security Officers (CISOs) evaluate risk severity.
According to Sakshi Grover, senior research manager for IDC Asia Pacific Cybersecurity Services, CISOs should assess three critical factors to determine material risks: first, which agents interact with untrusted content such as pull requests and issues; second, whether those agents’ outputs can directly or indirectly trigger higher-privilege workflows or agents; and third, the full capabilities of the involved identities, credentials, and tools.
Mapping Transitive Authority
Existing security solutions often provide only a fragmented view of authority transfer among agents and workflows. Grover noted that tools like IAM, PAM, CIEM, and application security might highlight individual identities and permissions but may fail to reveal their connection within a single event-driven delegation framework.
“Inventory lists what exists, while delegation mapping shows what could transpire,” Gogia articulated. Security teams should trace external inputs from their entry point to the downstream systems reacting to the outputs. This review must recognize the implications of shared workflow states, such as comments triggering specific commands.
“The pivotal question isn’t just if Agent A can call Agent B; it’s whether Agent A can manipulate what Agent B already trusts,” he added. Importantly, human intervention doesn’t always bridge this gap; for example, despite requiring a maintainer’s approval to merge a pull request, automation can still sway the evidence they see.
“An attacker does not need merge rights if they can manufacture compelling evidence for another to merge,” Gogia stated. He proposed that reviews should tie independently authenticated reviewers to the exact code scrutinized, and any notable changes should invalidate prior approvals. Grover further recommended treating alterations to reviews and approvals as distinct security events, which should be sent to a logging system unchangeable by the workflow’s own identity.