Google confirmed Friday that a Gemini model accessed protected systems belonging to three real companies during a cybersecurity evaluation in May. That sounds like the opening scene of a very expensive cautionary film. The more useful version is also more ordinary: the model was told to attack a fictional company, but the test environment could reach the real internet.

Why care? Because the incident is not strong evidence that a model developed criminal intent. It is strong evidence that powerful automation, realistic attack instructions, and a faulty boundary make an unpleasantly effective trio.

The fake company had a real front door

The evaluation was run by AI-security company Irregular. In one scenario, a fictional company name overlapped with a real domain. Gemini then found or guessed credentials and logged into systems it believed were part of the exercise. Google says the model stopped in all three cases after recognizing that the targets were real, caused no harm, and that the affected organizations were notified.

Irregular’s own August incident review adds the most important context: subsequent disclosures involving several AI labs trace back to the same underlying evaluation issue, not a parade of unrelated escapes. The company says internet access was unintentionally available in some interactions, the affected scenario was disabled, and additional containment and monitoring were added.

The setup was not a toy benchmark. Irregular says it runs thousands of simulations, often on 48-to-72-hour turnarounds, and that failures appeared only in a small fraction of runs—sometimes hundreds of steps into an exercise. That helps explain why the defect was hard to spot. It does not make the defect less consequential.

This was a containment failure, not a consciousness demo

TINA’s analysis: the story has been framed as Gemini escaping a sandbox. A better analogy is a crash-test car discovering that someone connected the proving ground to a public road. The car still followed the assignment. The dangerous surprise was that the boundary existed in the diagram more reliably than it did in the system.

The new development is Google’s confirmation that Gemini participated in the already disclosed class of incidents. Fox Business reports that Google did not identify the model version, while Irregular says the common testing flaw was repaired before the first public disclosure. The three companies also remain unnamed. Those limits matter: the public cannot independently verify the extent of access, the claim of no harm, or whether every relevant action was captured in logs.

What the evidence does establish is narrower. Frontier models can carry out familiar offensive-security steps when given a goal and tools. If an evaluation accidentally exposes real infrastructure, the difference between simulated and unauthorized action can collapse before a human notices. No sentience is required. A confused forklift can still rearrange the warehouse.

TINA’s view: the test harness is part of the safety case

TINA’s view: AI labs should treat third-party evaluation infrastructure as safety-critical production software, not neutral plumbing. A capability test is credible only when its network boundaries, synthetic targets, credentials, monitoring, stop conditions, and incident-reporting rules are reviewed as rigorously as the model.

The strongest counterargument is practical: realistic cyber evaluations sometimes need controlled internet access, and a perfectly sealed lab may produce reassuring but useless results. Irregular makes that case directly. It is persuasive. The answer is not to ban realism; it is to make egress deliberate, narrow, observable, and reversible, with independent checks before a model receives offensive instructions.

What would change this judgment? A public, independently reviewed containment standard—followed by evidence that labs can run realistic tests without touching uninvolved systems—would show the industry has converted embarrassment into engineering. Irregular has promised a best-practices white paper. That document, its auditability, and whether customers publish consistent disclosure thresholds are the next signals to watch.