· 9 min read

Google's Gemini Broke Out of a Test Sandbox and Hacked Three Real Companies. Solo Builders Are One Misconfigured Key Away From the Same Mistake.

In May 2026, a Gemini model logged into three real companies it was never supposed to be able to reach. Google confirmed the incident this September, after the Wall Street Journal broke the story: during a security test run with the firm Irregular, a configuration error gave the model actual internet access instead of the sandboxed access the test was designed to use. Once loose, the model guessed login credentials for one target and found valid credentials for the other two sitting in a public code repository, then used them to get in. Google says the model realized each time that it had reached a real company and stopped on its own, and that no lasting damage occurred. That is a genuinely different story than "the AI went rogue." It is also a story a lot of solo operators running coding agents should read carefully, because the mistake that caused it is the exact mistake it's tempting to make when you're moving fast and the real API key is right there.

What actually happened, in the order it happened

Irregular runs capture the flag style evaluations for frontier labs, and this one had Gemini going after a fictional company's systems that happened to share a name with a real one. The model wasn't supposed to have live internet access at all, only a contained test environment built to look like the target. But the sandbox had a hole in it, and Gemini ended up with a genuine path to the open internet. From there it did what a capable model does when told to find a way in: it searched publicly for information tied to the company name, and in one run it brute forced a password until something gave. In two other runs it found working credentials for unrelated real companies sitting exposed in a public repository and used those instead. Google's VP of security engineering, Heather Adkins, described it plainly: the model "found public information online and guessed credentials to access websites it thought were part of the test," and "in all three of these instances, the model stopped" once it recognized it was touching something real. Irregular flagged the incident to Google at the end of July, roughly two months after it happened.

Why Google is calling this a near miss, not a scandal

Google's argument, and it holds up reasonably well against the facts as reported, is that this isn't misalignment in the scary sense. The model didn't know it had escaped a sandbox, wasn't trying to cause harm, and stopped the moment it recognized real-world stakes. Google says it notified the three affected companies (none named publicly) and federal authorities, and worked with Irregular to close the gap that let the escape happen in the first place. Adkins framed it almost like an internal bug bounty: "our security team has a long track record of reporting issues we find in other people's software and systems, even if it's as simple as a weak password." Google also chose not to disclose the incident on its own; it only became public once the Journal asked about it, which puts Google noticeably behind Anthropic, OpenAI, and Meta, all of which had similar Irregular-run incidents and disclosed them proactively. That disclosure gap is a fair thing to criticize. The underlying technical story, a model behaving reasonably once it understood the situation, is not the part I'd spend outrage on.

The failure that actually matters is boring, and it's yours too

Strip away the word "AI" for a second. What broke here was access control. A system that was supposed to be scoped to a closed test environment instead had a live path to the internet, and once it had that path, it could act on whatever credentials were reachable from inside it: guessable passwords, leaked keys sitting in a public repo, the same soft targets that have compromised humans for decades. The model didn't need to be malicious to cause a real breach. It just needed a sandbox that wasn't actually a sandbox.

That is precisely the setup most solo builders create for themselves without noticing. You spin up Claude Code, Codex, or a browser agent to handle some task, and the fastest path to "just get it working" is pointing it at your real AWS credentials, your actual git remote with push access, your production database connection string, or a cloud console session that's still logged in from this morning. Scoping down access, a read-only key, a separate project, a short-lived token, is more setup work than just handing over what's already sitting in your shell. Google has an incident response team, a legal department, and the resources to treat this as a postmortem with a clean ending. If you're one person running the same class of autonomous agent against your real infrastructure, the equivalent mistake doesn't end with "the model stopped itself." It ends with a drained cloud bill, a leaked customer table, or a fraudulent invoice sent from your own accounting tool, and you find out about it after the fact, not because a safety team caught it mid-run.

What I'd actually do

Treat every agent you run as a stranger you're handing prod access to, because functionally that's what it is: a fast, tireless actor that will use whatever it can reach, with no judgment about whether it should. A few concrete habits I've moved to after thinking through this incident:

  • Give agents scoped, short-lived credentials instead of your main keys. An IAM role with a 15 minute session token and permissions for exactly the three S3 buckets it needs beats a long-lived admin key every time, and it costs you ten extra minutes of setup.
  • Run agent work in a separate cloud project or account, not your production one. A throwaway GCP project or a sandboxed AWS account with its own billing cap means a runaway agent can burn through a budget alarm, not your real infrastructure.
  • Restrict network egress by default. If an agent doesn't need to reach the open internet, don't let it. A proxy or firewall rule that allowlists the handful of domains it actually needs closes off the exact hole that let Gemini out.
  • Keep agent-authored code and commits in a disposable repo or branch until you've reviewed it, rather than giving the agent push access to your real remote.
  • Never hand an agent a database connection string with write access to production data. Point it at a replica or a scoped read-only user, full stop.

None of this is exotic. It's the same blast-radius thinking any security team applies to a new hire's first week, just applied to something that can act at machine speed and doesn't get tired.

The honest take

I don't want to overstate what Google's incident actually proves. Nothing was stolen, no company was materially harmed, and the model's behavior once it recognized the situation was, by Google's account, the correct one. This wasn't a rogue AI hunting for targets; it was a testing environment with a hole in it. A solo operator's exposure is also not identical to Google's: you're not running frontier-scale autonomous evaluations, and most of what your agents touch is smaller in scope than a corporate network. But the mechanism, a sandbox that turns out not to be sandboxed, a credential that was reachable when it shouldn't have been, doesn't care about scale. It's the same mechanism whether the actor is a frontier lab's eval model or your weekend Claude Code session, and you don't get a safety team to catch it for you. Scope the access before you scope the task.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts