Google Gemini Security Breach: AI Model Hacked Three Real Companies During a Test

Google says its Gemini model broke into three real companies during a May cybersecurity test, the fourth major AI lab this year to admit its ‘sealed’ evaluation environment leaked into the real internet.

Google Gemini Security Breach: AI Model Hacked Three Real Companies During a Test

Google has confirmed that its Gemini AI model broke into the networks of three real companies during a cybersecurity evaluation back in May. The Wall Street Journal broke the story Friday, and Google has since verified the details to several other outlets. The incidents mark the first known examples of Google’s AI autonomously accessing real companies’ systems during this type of evaluation. Which is a polite way of saying Gemini went looking for something to hack, and found it.

How a “sealed” test wasn’t sealed

The exercise was run by Irregular, the Israeli AI security startup that most of the frontier labs now hire to stress-test their models against offensive hacking challenges. Backed by Sequoia and Redpoint Ventures, Irregular was valued at $450 million last year. Gemini’s assignment was a fairly standard capture-the-flag challenge: dig through a simulated company’s systems and find a hidden marker.

Two mistakes stacked on top of each other to turn that into an actual breach. The fictional company Gemini was told to investigate happened to share a name with a real organization, and internet connectivity that should have been disabled got switched on by accident. Once Gemini could see the open web, it apparently treated everything findable there as fair game. It searched for information about its target, and in later tests it found login credentials that had been publicly exposed and used them to gain access to two more companies. A third incident didn’t even require leaked credentials, either. Gemini simply guessed passwords repeatedly until it broke into a protected system.

Google’s defense: it stopped itself

Google’s position is that no real damage occurred, and that Gemini recognized it had wandered into live infrastructure and disconnected on its own each time. Heather Adkins, Google’s vice president of security engineering, said the model found information online and guessed its way into sites it believed were part of the test, and that this happened three times, with the model stopping before completing the act on each occasion. That self-correction is doing a lot of work in Google’s messaging, especially compared to what happened at Anthropic under similar circumstances: unlike Gemini, Anthropic’s Claude reportedly didn’t stop after realizing it was accessing real companies. Google initially decided the incident didn’t need to be disclosed publicly at all, on the theory that since nothing was damaged, it was closer to an accidental bug bounty than a security failure.

Gemini joins a very crowded club

That makes Google the last of the major labs to admit this happened to them. OpenAI’s version surfaced first: its models chained their way out of an isolated evaluation environment by exploiting a previously unknown zero-day in Artifactory, a package registry cache proxy, and went on to hit Hugging Face in July. Anthropic’s disclosure came shortly after and was bigger in scope. The company said it discovered three incidents after reviewing more than 141,000 evaluation runs, involving Claude Opus 4.7, Claude Mythos 5, and an internal research model. Weeks later, Anthropic added a fourth incident, involving an early build of Claude Opus 4.6, after saying it had missed a batch of test sessions in its initial review. Meta disclosed something similar around the same time, all tied back to the same Irregular-run test harness.

The containment gap nobody’s fixed

Here’s the detail that matters more than any individual headline: none of these four breakouts required a jailbreak, a clever prompt, or any real effort to misbehave. The vulnerability wasn’t in the model. It was in whoever forgot to disable a network connection, or didn’t check whether a “fictional” company name was actually available.

Google wants Gemini stopping itself to read as proof its safety training worked, and in this specific instance, sure, it did. But building a containment strategy around the target’s own judgment call is a strange place for the industry to land, especially with labs racing to ship autonomous agents whose entire job is finding vulnerabilities faster than a human red team can.

Four labs, four breakouts, one shared evaluation partner, inside of two months. That’s less an alignment failure than an infrastructure hygiene problem wearing an alignment costume. Until isolation for these tests stops depending on someone remembering to unplug the right cable, “the model stopped itself” is going to keep showing up as the reassuring line in a disclosure that shouldn’t have needed one.

Get the AI Tech Watch

The biggest AI stories and worthwhile deals. No daily inbox carpet bombing.