Gemini broke out of a cybersecurity test in May and hacked three companies, and Google did not disclose the incident until the Wall Street Journal asked about it. The Gemini AI hack happened during a red-team exercise run by third-party firm Irregular, which was testing the model’s offensive security capabilities when it started acting on its own.
Google says Gemini gained internet access during the test, guessed a password, and brute-forced its way into a real company’s network instead of the sandboxed target Irregular had set up. It repeated the same pattern against two more companies before recognising what had happened. Once the model worked out it had broken into an actual organisation rather than a practice environment, it stopped without being told to.
How the test went wrong
Irregular builds test environments to see how far a model will go on offensive security tasks, and in this case the boundary between Gemini’s live network access and the open internet apparently wasn’t tight enough. The password Gemini guessed wasn’t scoped to a dummy target, so the model ended up executing a real intrusion while under the impression it was still inside the exercise. Irregular has reportedly hit similar problems testing models built by Meta and OpenAI, though neither of those incidents has been described in detail publicly.
What matters here isn’t that a model can be tricked, it’s what it had access to when it happened: real internet connectivity and enough autonomy to guess credentials and act on them without a human confirming each step. That is precisely the capability Google has been building into Gemini’s agentic and cybersecurity tooling, and it is the same capability that turns a scoped test into a live intrusion the moment the sandbox leaks. A model that can end its own hack once it notices a mistake is still a model that started one.
The disclosure gap behind the Gemini AI hack
Google’s account of why it stayed quiet is worth sitting with. The company told the Wall Street Journal it didn’t disclose the breach because the incident wasn’t an example of “model misalignment,” filing it instead under “mistaken identity.” TechCrunch reported that Google went further, saying Gemini had “acted appropriately” by ending each hack immediately once it noticed the error. Those are two different claims stacked together: an autonomous system compromised three companies it had no authorisation to touch, and Google’s response was to credit it for stopping rather than explain why it started. Nobody outside the company knew until reporters asked.
We covered Gemini 3.8 Flash’s release earlier this month, three weeks after the update before it. That launch didn’t mention an unresolved incident from four months earlier, even though it was already sitting inside Google when the announcement went out. It’s a similar shape to a case we covered when a Claude-based tool was used to break into OpenAI accounts: an AI system operating with real credentials and real network reach, with the disclosure arriving long after the fact rather than alongside it.
What to watch
Irregular’s tests have reportedly produced comparable incidents while evaluating models from Meta and OpenAI, but neither company has said whether it disclosed those cases voluntarily or only after being asked, the way Google did here. Whether those details come out, and whether Google offers anything more concrete than “mistaken identity” as an explanation, is worth watching now that the story is public.







