Who Checks The AI Sandbox Before The Agent Goes In?

In May, Google's Gemini broke into three real companies during a security test and stopped itself each time. The model held; the AI sandbox did not. What that means for developers running evaluations, and for any organisation piloting agents with real credentials.
AI generated image - an AI sandbox test rack in a wire cage with its door open and a network cable running out to live ports

In May, Google’s Gemini broke into three real companies during a security test. Each time, it stopped once it realised the targets were real.

That detail is the reassuring one, and it is the wrong one to take home. The model’s judgement held. The AI sandbox around it did not, and the AI sandbox is the part your organisation builds for itself.

What happened inside the AI sandbox

The test was run by Irregular, an independent company that conducts cybersecurity evaluations. During a standard evaluation, the environment gave Gemini internet access it was never meant to have, as Al Jazeera and Reuters reported. The model then found public information online and guessed credentials for three websites it believed were part of the exercise.

In one case it guessed passwords until it reached a protected system. In the other two, it found working credentials in a public repository, as the Wall Street Journal first reported. Google says the model ceased in all three instances. It also says the three organisations were told and that its testing partner has since changed its processes.

Irregular told Google at the end of July. Google says the behaviour was not misalignment and did not warrant public disclosure, because the safety measures worked.

Nor is this a one-off. Meta, Anthropic and OpenAI have disclosed similar incidents linked to the same evaluator.

The developer’s side of the AI sandbox

For the developer, testing of this kind is not optional housekeeping. Article 55(1)(a) of the AI Act requires providers of general-purpose AI models with systemic risk to evaluate their models, including conducting and documenting adversarial testing. Article 55(1)(d) adds a duty to ensure adequate cybersecurity protection for the model and its physical infrastructure.

Read those two duties side by side and the tension becomes plain. Adversarial testing of cyber capability means pointing a capable model at targets. The protection duty means making sure the targets are the intended ones. Here the AI sandbox was the fence. In May, the evaluation worked and the fence did not.

When a third party builds the fence

There is a value-chain lesson here too. The environment belonged to the evaluator, and so did the fix. Google’s answer was to work with its training partner on changes to that partner’s testing processes.

Outsourcing an evaluation does not outsource its consequences. Three companies with no part in the test had systems accessed all the same. Whoever writes the evaluation contract therefore has to decide who checks the harness before the model goes in.

Why deployers inherit the same problem

A deployer is not red-teaming frontier models. It is, however, running pilots. An agent receives a task, a set of credentials, a network connection and a deadline, usually inside something the project team calls an AI sandbox.

The phrase in Google’s statement worth rereading is a small one. The model accessed websites it thought were within the scope of its test. In other words, scope lived in the model’s understanding of the task. Nothing in the environment enforced it.

That is the design choice worth checking in every AI sandbox your organisation runs. An agent with good intentions and the wrong credentials still holds the wrong credentials.

Four places an AI sandbox needs its scope

Scope that exists only in a prompt is a request, not a boundary. In a pilot, it belongs in four places:

  • The network. An AI sandbox with open internet access is a test on the live internet with extra steps. An allow-list of what the agent may reach, with everything else blocked by default, turns scope into something the agent cannot misread.
  • The credentials. Test accounts, scoped to the pilot and expiring when it ends. Anything real that the agent can find, it can use. Two of the three Gemini cases began with credentials sitting in a public repository.
  • The data. Synthetic or copied data wherever the task allows. A boundary failure then exposes nothing that belongs to a customer or an employee.
  • The stop. A named person with the authority and the technical means to halt the run. Alongside that, a record of everything the agent touched before it was halted.

None of this is exotic. It is the same containment a security team would expect for a penetration test, applied to a tool that nobody thought of as a tester.

Who owns the fence in your organisation

Think about who set up your last agent pilot. If the answer is an innovation team, a business unit or the vendor’s own solutions engineer, security may never have seen the AI sandbox at all.

This is where training earns its keep. The team running the pilot needs to treat “AI sandbox” as a claim to be checked, not a label to be trusted. Meanwhile, the AI lead needs a short way to check it before the first run, and the standing to delay the run if the answer is unclear.

The question worth asking before the next pilot

The Gemini episode ended quietly because a model recognised a real company and stopped. That is a welcome finding about the model. It is an uncomfortable one about the AI sandbox.

A control that depends on the agent noticing its own mistake is not one you would accept from a human contractor. Consequently, the useful question for your vendor and your own team is not whether the model behaves well. It is where the boundary is enforced, and who confirmed it was closed before the agent started.

If your organisation already keeps a log for AI behaviour below the serious incident line, an agent reaching a system outside its pilot belongs in it, even when the agent stops by itself.

Asking that question well is a governance skill, and it can be taught. The AIGP course on Future Prep Applied is built for the people who have to ask it inside an organisation, and for the AI lead who then has to bring the pilot team along.

Bas Hennis

Future Prep helps organizations prepare for the impact of AI and emerging technologies. We provide hands-on training, strategic advice, and smart tools for the responsible use of AI, governance, and digital resilience.LinkedIn

Newsletter
Related Blogs
LATEST NEWS

AI governance is not a future problem

Regulation is already in effect. Your competitors are already building internal capability. The gap between ‘we are aware of AI’ and ‘we have operational control’ is closing, and it closes faster with a structured framework.

 

Book a 30-minute discovery call. No obligation. We will assess where your organisation stands and what a realistic starting point looks like.

No sales pressure. No jargon. Just a structured conversation about your organisation's AI readiness.

Scroll to Top