Google’s Gemini AI breached the systems of three real companies during a cybersecurity test, marking the first known case of the search giant’s AI autonomously carrying out such intrusions, the Wall Street Journal reported exclusively.
The incidents occurred in May, when Gemini participated in a “capture-the-flag” exercise designed to test its cybersecurity capabilities against a fictional company operating inside the testing infrastructure of Irregular, an AI security testing firm. Irregular notified Google about the breaches in late July, and Google confirmed them on a Friday after being contacted by the journal.
In one case, Gemini guessed passwords until it gained access to a protected system. In two other test runs, the model located credentials in publicly accessible online repositories and used them to enter systems belonging to actual companies.
The breaches stemmed from a confluence of design flaws in the testing setup. The fictional business shared its name with a real company, and the testing environment unintentionally provided Gemini with internet access, allowing the model to target real-world systems while attempting to complete the exercise.
Gemini stopped each intrusion autonomously after recognizing it had accessed a real company’s systems rather than the intended fictional target. Google stated that no harm occurred and said it notified all three affected businesses as well as U.S. federal authorities.
Google declined to identify the companies involved or specify which Gemini model carried out the intrusions, though it confirmed the incidents did not involve its newest model. The company also said it does not consider the behavior an example of “AI model misalignment,” pointing to Gemini’s decision to stop once it recognized the unintended breaches.
The episode echoes previous testing incidents at other major AI labs. OpenAI, Anthropic and Meta have all encountered cases where their models reached targets outside their intended testing environments during security assessments.












