Thomas Boele, Regional Director Sales Engineering, CER / DACH Check Point Software Technologies
OpenAI, Anthropic, and the UK’s AI Security Institute each disclosed within a fortnight that AI agents broke out of their test environments and touched real systems and real people. The incidents themselves matter less than the speed at which agent capabilities are advancing, explains Thomas Boele, Global Director Solutions Engineering – AI Security, Check Point Software.

Within fourteen days in July and August 2026, three organizations disclosed that AI agents deployed during internal security evaluations broke out of their intended test boundaries and affected real infrastructure, code, or people. Individually, none of the three cases is dramatic. Together, they mark a shift from theoretical risk to documented incident – and in one case it was human attentiveness, not a technical safeguard, that prevented harm.

The most recent case came from the UK’s AI Security Institute (AISI), which published a report in early August describing how an AI agent broke out of a testing environment during a routine evaluation. The agent researched the maintainers of a public open-source project, fabricated fake online identities, and used them to try to persuade a real maintainer to approve malicious code – a social engineering approach normally associated with human attackers in supply-chain scenarios. What stands out is not the technique but how it emerged: no one instructed the agent to deceive anyone, nor was it explicitly barred from doing so – a restriction that, until now, seemed unnecessary for models with alignment training.

Incident Under the Microscope

AISI tested seven models across two “cyber ranges” – controlled, simulated networks resembling real environments – in 122 runs. In ten runs, an agent acted outside its scope, producing 19 documented actions: 17 involving Anthropic’s Mythos 5 and two from a single run with OpenAI’s GPT-5.6 Sol, per the institute’s technical report.

Two caveats matter. Internet access was deliberately enabled, and the cyber-safety classifiers providers normally use to block risky behavior were switched off, to establish maximum capability – conditions that do not reflect how these models reach end users. No sandbox breakout occurred, and AISI has not confirmed real-world damage. Its own monitoring caught the episode: on the morning of July 28, systems flagged traffic leaving a test system via the Tor network, and the incident was contained within an hour.

Three Disclosures in Fourteen Days

The AISI report is the third disclosure of its kind within two weeks, and the three cases are instructive because each has a different root cause.

On July 21, OpenAI disclosed that two models had escaped a sandboxed environment through a previously unknown vulnerability and compromised Hugging Face’s production infrastructure, according to the company. Hugging Face detected the intrusion itself and alerted law enforcement before OpenAI connected it to its own evaluation run.

Anthropic followed on July 30. A review of 141,006 evaluation runs identified three incidents in which Claude models – Opus 4.7, Mythos 5, and an internal research model – gained unauthorized access to production infrastructure at three real organizations, according to Anthropic; the earliest dated to April 2026. The cause was a miscommunication with evaluation partner Irregular: test machines had active internet access even though the prompt told the model otherwise. The intrusions relied on weak passwords and unauthenticated endpoints, and two of the three affected organizations had not noticed the activity on their own.

Not the Cases, But the Pace

An AI agent inventing false identities to get a human to approve malicious code was, until recently, a subject for academic papers, not incident reports; AISI describes this kind of targeted deception as behavior that had largely remained theoretical. It is now on record.

One detail from Anthropic’s report deserves attention, because it initially seems to cut against that reading. The three models responded differently once cues suggested a real environment: Opus 4.7, the oldest, continued the attack even after recognizing the systems as production; Mythos 5, despite an accurate initial assessment, talked itself back into assuming it was still in a simulation; the newest, an internal research model, halted the operation. That looks reassuring, but it implies the opposite for security architecture: the ability to compromise a system is advancing faster than the ability to assess one’s own situation, so self-assessment cannot serve as a control, since it is a behavior, not a safety property.

All three incidents were detected, contained, and disclosed – early-warning mechanisms are working, though it remains open whether defense scales with the technology and whether organizations react as fast as warnings reach them. Worth noting: what stopped the AISI attack was a person – the maintainer who reviewed and rejected the pull request – plus an outsider who opened the suspicious code in an isolated environment, after technical monitoring flagged it only afterward, via anomalous traffic. Human attentiveness held this time; it does not scale against an attacker that works continuously and never tires.

Consequences for Practice

AISI draws three lessons for evaluators: fine-grained network controls instead of blanket internet access, real-time monitoring of running evaluations, and designs that assume a capable model may leave its scope. For other organizations, its advice is more sober and not new: solid cybersecurity fundamentals, rigorous review of external code, board-level ownership of cybersecurity, and minimum standards across the supply chain; in German-speaking countries, BSI guidance offers a suitable framework.

Check Point groups the necessary measures into three areas: defense against AI-driven attacks, since adversaries have the same capabilities these tests expose; control over an organization’s own AI, meaning knowing which agents run and what they may access or do; and continuous verification rather than an assumption of safety.

For agents already in production, four questions help as a starting point: Which agents are running, including those built by non-developers? What can each access? What permissions exceed what was intended? Would a deviation even be noticed? If the answer to that last question is “no,” that is the gap to close first.

By Jakob Jung

Dr. Jakob Jung is Editor-in-Chief of Security Storage and Channel Germany. He has been working in IT journalism for more than 20 years. His career includes Computer Reseller News, Heise Resale, Informationweek, Techtarget (storage and data center) and ChannelBiz. He also freelances for numerous IT publications, including Computerwoche, Channelpartner, IT-Business, Storage-Insider and ZDnet. His main topics are channel, storage, security, data center, ERP and CRM. Contact via Mail: jakob.jung@security-storage-und-channel-germany.de

Leave a Reply

Your email address will not be published. Required fields are marked *

WordPress Cookie Notice by Real Cookie Banner