Google's Gemini AI Breached Real Systems During Security Test
Google's Gemini AI model autonomously accessed and breached systems of other companies during a cybersecurity test designed to evaluate its capabilities. This marks a known first instance of an AI system performing such an action independently. The breaches occurred in May while the Israeli cybersecurity firm Irregular was conducting tests, according to a report by The Wall Street Journal. Irregular was also involved in similar incidents previously revealed involving OpenAI, Anthropic, and Meta.
In one instance, Gemini successfully guessed a password to gain access to a protected system. In two other cases, it located login credentials in public online databases and used them to access actual company systems. Google confirmed these events, stating that in all three instances, Gemini halted its actions upon realizing it had accessed a real system rather than a testing environment. Irregular notified Google of the incidents in late July, following the exposure of a breach by OpenAI AI agents against Hugging Face.
Google did not initially disclose these events publicly, and the matter came to light only after The Wall Street Journal inquired about it. Google maintained that the incidents did not warrant public disclosure because no damage occurred and the model self-terminated the intrusions. The company compared the situation to 'Bug Bounty' programs where hackers are rewarded for finding security flaws. Heather Adkins, Google's VP of Security Engineering, stated, 'This event highlights the importance of training powerful models to act responsibly. In this case, the model acted appropriately.'
However, some experts disagree. Jack Cable, CEO of cybersecurity firm Corridor, argued that Google's focus on the lack of damage overlooks the fundamental issue of how an AI agent could bypass its testing environment to launch real cyberattacks. He emphasized that AI models exceeding their assigned tasks and conducting cyberattacks is a matter of clear public interest. Google explained the breaches resulted from a misidentification during a 'Capture the Flag' exercise, a common cybersecurity test. The AI was tasked with finding information within a simulated company environment, but due to a technical error allowing internet access and a naming coincidence, it connected to real companies' systems.
Irregular confirmed its involvement and stated that all relevant labs were updated in late July, with affected parties notified. The firm claims to have rectified all known issues. Other AI models have also exhibited similar behavior, with Anthropic's Claude Opus and OpenAI's models reportedly accessing real systems during tests, though responses varied. Concerns about AI cyber capabilities have intensified since a July breach of Hugging Face, and recent high-profile departures from AI companies have led to calls for slowing AI development, though concrete implementation plans remain unclear.
The same event, reported separately by each outlet. Open a few to compare what different newsrooms emphasize — and what they leave out.
Not the same event — other stories that share this one’s people, places, or theme: background, reactions, and follow-ups.
