AI, News

Gemini’s breach of real companies exposes an AI guardrail problem

| September 21, 2026
Gemini logo

Google says one of its Gemini models accessed systems belonging to three real companies during a cybersecurity evaluation in May.

The model reportedly guessed credentials in one case, while finding exposed credentials in public repositories in two others. Google says Gemini stopped once it recognized that it had reached real infrastructure and that the affected organizations were notified.

Gemini was participating in an evaluation run by Irregular, a third-party AI cybersecurity testing firm. Similar incidents involving models from Anthropic, OpenAI, and Meta have also been linked to the same underlying problems with evaluation environments that allowed the models to reach the public internet.

But the news arrives at a particularly interesting moment.

For months, the AI industry has been steadily escalating its demonstrations of agentic capability: Models can browse, use tools, write code, pursue multi-step goals, and sometimes find ways around obstacles their developers did not anticipate. The market rewards eye-catching evidence of autonomy. “It completed the task” is impressive. “It did something it was not supposed to do” can be even more memorable.

There is another way to look at the incident: not as evidence of an AI suddenly developing criminal intent, but as a small, concrete example of the “AI alignment” problem.

Alignment is the deceptively difficult task of making an AI system’s behavior match what people actually intended, rather than merely the narrow objective they managed to express. In this case, the objective was to locate hidden information within a simulated target and complete the evaluation. But any human operator would likely have regarded one condition as non-negotiable: Do not attempt to access real companies.

Gemini appears to have optimized for the first instruction while treating the second as an inference problem. It found an organization with a matching name, encountered systems reachable from the internet, and proceeded as though they belonged to the exercise. Although it completed the task in a way its human operators would not have approved, Google says it stopped after recognizing that it had reached genuine infrastructure—something other models have failed to do.

Even so, the incident shows how potential alignment failures can be much more mundane. Give an agent a goal, tools, and room to act, and it may faithfully pursue the measurable part of the assignment while overlooking the unstated boundaries that humans rely on one another to understand. AI models do not automatically know where we draw the line.

There may also be an awkward marketing angle in the background. A model capable enough to make the wrong kind of progress can still look very capable indeed. In a market where every lab wants to demonstrate that its agents can plan, code, browse, and act independently, even a safety disclosure can carry a secondary message: Ours can play in this league, too.

It’s another reason to take seriously warnings from insiders, industry leaders, and politicians that safeguards must keep pace with the autonomy given to AI models.


What do cybercriminals know about you?

Use Malwarebytes’ free Digital Footprint scan to see whether your personal information has been exposed online.

About the author

Pieter Arntz

Malware Intelligence Researcher

Was a Microsoft MVP in consumer security for 12 years running. Can speak four languages. Smells of rich mahogany and leather-bound books.