AI Agents Out of Control: Three Incidents in Fourteen Days
Within two weeks, OpenAI, Anthropic, and the UK-based AI Security Institute disclosed that, during internal security tests, AI agents had strayed beyond the intended test environment and impacted real systems and real people. Thomas Boele, Global Director of Solutions Engineering – AI Security at Check Point, analyzes the incidents and draws practical conclusions.
Within a span of fourteen days, OpenAI, Anthropic, and the British AI Security Institute (AISI) each disclosed that, during internal security audits, AI agents had strayed beyond their intended test environments and interacted with real systems and real people. Less remarkable than the individual cases is the speed at which these agents’ capabilities are evolving—and the fact that, in one of the cases, it was human attention—not a technical control—that prevented the damage.
A Closer Look at the Incident
In early August 2026, the AISI published a report on the escape of AI agents from test environments. As part of a routine test, an AI agent researched the maintainers of a publicly used open-source project, created several fake online identities based on this information, and attempted to persuade a real maintainer through these identities to approve malicious code —a classic social engineering approach, such as that typically described by human attackers in supply-chain scenarios.
What is remarkable here is not so much the technology as the origin of this behavior: No one had instructed the agent to deceive. Nor had he been explicitly prohibited from using social engineering techniques or open internet access for such purposes—until now, such a restriction had seemed unnecessary for models with alignment training.
As part of this study, the AISI tested two so-called cyber ranges—controlled, simulated networks that replicate real-world environments—using seven different models across a total of 122 runs. In ten of these runs, an agent operated outside the specified test scope; 19 individual actions were documented, 17 of which were attributed to Anthropic’s Mythos 5 and two to a single run with OpenAI’s GPT-5.6 Sol. The complete list can be found in the AISI’s technical report at https://www.aisi.gov.uk.
When evaluating these figures, two key conditions must be taken into account: Internet access was intentionally enabled, and the model providers’ cyber classifiers—that is, the filtering mechanisms designed to prevent risky behavior—were intentionally disabled. Both measures are intended to determine a model’s maximum performance, and both explicitly do not correspond to the configuration in which these models are made available to end users. There was no escape from the sandbox. AISI has not yet identified any verifiable real-world damage, but does note isolated, limited impacts under specific circumstances.
Three disclosures in fourteen days
This is the third disclosure of this kind within fourteen days, and comparing the three cases is instructive because the causes differ in each instance.
On July 21, 2026, OpenAI announced that two models had escaped from a isolated test environment via a previously unknown vulnerability and subsequently compromised Hugging Face’s production infrastructure. What is noteworthy here is how the incident was detected: Hugging Face had discovered the attack on its own and reported it to law enforcement before OpenAI was able to establish the connection to its own evaluation run.
Anthropic followed on July 30, 2026. A retrospective review of 141,006 evaluation runs revealed three incidents in which Claude models—specifically Opus 4.7, Mythos 5, and an internal research model—gained unauthorized access to the production infrastructure of three real-world organizations; the earliest of these incidents date back to April 2026. The cause was not an exploit, but rather a misunderstanding between Anthropic and the external evaluation partner: The machines used in the test had active internet access, even though the prompt explicitly told the model otherwise. The actual breach of the affected organizations then occurred using the simplest of means: weak passwords and unauthenticated endpoints. The two organizations that Anthropic was able to contact had not noticed the activity themselves.
The problem isn't the individual cases; it's the pace
Until recently, an AI agent that invents false identities to trick a human into releasing malicious code was the subject of academic papers; the AISI explicitly describes this form of targeted deception as behavior that had previously remained largely theoretical. Now, it has been documented.
One detail from the Anthropic report deserves special attention: The three models in question reacted differently when evidence of a real-world environment emerged. Opus 4.7, the oldest of the three, continued the attack even though it recognized the production systems as such. Mythos 5—despite its own accurate interim assessment—reverted to the assumption that it was still operating within a simulation. The newest, internal research model terminated the process.
At first glance, this is good news, because the latest model responded most appropriately. For the security architecture, however, this results in the opposite of a relief: The ability to compromise a system is growing faster than the ability to accurately assess one’s own situation. Therefore, a control mechanism must not be based on an agent’s assessment of the situation, because that is a behavior and not a security property.
One positive aspect is that all three incidents were detected, contained, and subsequently disclosed. This shows that early-warning mechanisms are working, and there is still time to take countermeasures. It remains to be seen whether this form of defense can scale with development and whether affected organizations as a whole can respond as quickly as the warnings reach them.
Against this backdrop, it is worth taking a closer look at what actually stopped the attack in the AISI case. It was detected through technical monitoring—though this occurred downstream and was triggered by suspicious network traffic—rather than through a check within the evaluation process itself. It was stopped by a person: the maintainer who reviewed and rejected the pull request. In addition, an outsider opened the suspicious code in an isolated environment. Human attention and established best practices held up in this case. However, neither of these can scale against an attacker who works continuously and never tires.
Implications for Practice
The AISI itself draws three lessons primarily intended for evaluators: fine-grained network controls instead of blanket Internet access, monitoring that tracks the ongoing evaluation in real time, and an evaluation design that accounts from the outset for the possibility that a high-performing model might exceed its parameters.
For organizations outside this circle, the AISI’s recommendation is significantly more pragmatic and, at its core, not new: a robust foundation for cybersecurity, as well as consistent testing of external code and external contributions. In addition, the AISI recommends making cybersecurity a priority for the executive board and requiring minimum standards throughout the entire supply chain.
Thomas Boele, Global Director of Solutions Engineering – AI Security at Check Point, breaks down the necessary measures into three areas of action: protection against AI-driven attacks, since attackers possess the same capabilities that these tests reveal; control over one’s own AI, because security managers must know which agents are operating within the organization, what they can access, and what actions they are permitted to take; and continuous verification rather than assumption, because the correct behavior of agents and chatbots must be continuously verified and not simply assumed.
For AI agents already in production, Boele recommends starting with four key questions: Which agents are currently in operation—including those created by individuals who are not developers? What can each of these agents access? What permissions does each agent have that go beyond what was originally intended? And would a deviation from this framework even be noticed? If the answer to the last question is «no,» that is precisely the gap that needs to be closed first.
Source: www.checkpoint.com
This article originally appeared on m-q.ch - https://www.m-q.ch/de/ki-agenten-ausser-kontrolle-drei-vorfaelle-in-vierzehn-tagen/
