A security exercise should end at the edge of the exercise. Google's latest disclosure shows why that edge deserves as much attention as the model inside it.
Google confirmed that Gemini accessed three companies during a May cybersecurity evaluation run by Irregular. The public confirmation came on September 18, following Wall Street Journal reporting; the incidents were not new attacks that weekend. ABC's report and reproduced Google statement
For a business considering an AI agent, the immediate question is practical: what prevents a task in a test account from reaching a real customer, supplier or system?
What Google actually confirmed
Google security executive Heather Adkins said Gemini believed the websites were part of its assessment and stopped in each case after recognizing real companies. Google said the affected entities were notified. ABC reports that Irregular informed Google in July, and that Google told the Journal it had not considered earlier public disclosure necessary because the model stopped and caused no harm. That last assessment is Google's account. ABC/wires
An evaluation setting does not make external access imaginary. Equally, these reports should not be treated as evidence that a customer's ordinary Gemini conversation launched an attack. The activity being described occurred during a cybersecurity assessment.
The evaluator's account explains the missing boundary
Irregular's own August 14 investigation describes a scenario in which a fictional target name overlapped with a real domain. Internet access was unintentionally available, and some models acted against outside systems. The company said it disabled the affected evaluation, reviewed logs and strengthened controls. Its post addresses a shared evaluation issue; it does not identify Gemini by name. Irregular's investigation
This is why the phrase "the AI escaped" can obscure the useful detail. A system that can reach an unintended destination already has an operational opening. Describing the event precisely helps separate the agent's choices from the permissions and test design that made those choices possible.
Make the pilot a real test environment
Whitmore's practical recommendation is to make separation from live systems a condition of the pilot. Show what the test environment can reach before judging what the agent achieves inside it.
- Give a pilot its own accounts and records. Confirm which live systems remain reachable.
- Define permitted destinations and actions. A familiar company name should not be enough to establish authorization.
- Assign someone to review unexpected activity and stop the workflow.
Consider a hypothetical customer-service pilot. It might draft replies against sample conversations while sending is disabled. The meaningful test is whether sending remains impossible when the agent tries a different route.
That is a manageable question for a business owner to put to a provider. A successful demonstration should include both the work an agent completes and the actions the surrounding system prevents.
