Andy Smith, Certified Instructor SANS Institute
Following a recent incident involving Meta, reports are increasing of AI agents escaping test environments and compromising other IT systems. Andy Smith, a Certified Instructor at the SANS Institute, explains that the problem is not a lack of awareness, but that most security teams have never actually reviewed how such attacks behave inside their own environments.

According to the SANS Institute, another case became public this week in which an AI agent operated by Meta broke out of a test environment. The incident adds to a growing number of reports that AI agents from various vendors are escaping sandbox environments and compromising other software providers over the internet. In the assessment of Andy Smith, a Certified Instructor at the SANS Institute, the actual weakness does not lie in a lack of awareness within the security industry. Practitioners are well aware that AI-driven attacks are occurring, he said; they have read the headlines. The gap, according to Smith, is that most teams have never actually reviewed how such attacks behave inside their own environments — behavior that challenges many of the assumptions underlying existing detection programs.

“Both perspectives have merit,” Smith is quoted as saying. He said progress is clearly visible in AI’s ability to find and exploit vulnerabilities, while a large number of agents can also be deployed together for an offensive operation, which he described as a substantial efficiency gain for attackers. Even so, he noted, an AI still has to go through the same steps a human attacker would. It has been known for years, Smith said, that forgotten or misconfigured resources are often what gives attackers initial access to a company — a finding that, in his view, continues to hold in the age of AI. He also said it reflects poor practice that some companies running sandbox tests with AI agents have not implemented sufficiently robust controls, which has led to attacks affecting other parties.

Based on the publicly known incidents, the SANS Institute outlines several questions security teams should be asking themselves:

Severity scoring for alerts

Existing rules often assume an active attacker or a clearly identifiable critical event. An autonomous agent may present neither. Preparing for such incidents means having severity logic designed to detect slow activity spread across multiple paths, where no single event appears critical on its own, while still ensuring that on-call staff are notified even outside regular working hours.

Automation as cover

If security teams do not have a clear picture of how their own legitimate automation normally behaves, hostile automation can blend in seamlessly. Conversely, if a team’s own system behavior is well understood, unfamiliar behavior stands out more quickly.

Deception measures

According to the SANS Institute, deception technology tends to work better against agents than against human attackers: while a cautious human attacker might avoid triggering a canary token, an agent issuing thousands of commands could trip one within the first hour. One affected security team reportedly acknowledged that it should have deployed deception measures; according to the SANS Institute, most environments still lack them.

Self-hosted open-weight models

A fourth question, which the SANS Institute says is rarely asked so far, concerns models that are part of an organization’s incident-response contingency plan: have they been checked for hidden behavior? Since such a model becomes part of the security infrastructure, it warrants the same scrutiny as any other component teams rely on during an incident.

The exercise is intended primarily as a diagnostic. If the answers turn out as they are likely to for most teams, security teams will end up with a concrete, prioritized list of improvement measures grounded in a real incident rather than a hypothetical scenario. The SANS Institute has also made a readiness check available so that security teams can assess where they actually stand.

As further incidents come to light, pressure is likely to grow on security teams to review their detection and response processes specifically against the behavior of autonomous AI agents — particularly around severity logic, deception measures, and self-hosted models used in incident response.

By Jakob Jung

Dr. Jakob Jung is Editor-in-Chief of Security Storage and Channel Germany. He has been working in IT journalism for more than 20 years. His career includes Computer Reseller News, Heise Resale, Informationweek, Techtarget (storage and data center) and ChannelBiz. He also freelances for numerous IT publications, including Computerwoche, Channelpartner, IT-Business, Storage-Insider and ZDnet. His main topics are channel, storage, security, data center, ERP and CRM. Contact via Mail: jakob.jung@security-storage-und-channel-germany.de

Leave a Reply

Your email address will not be published. Required fields are marked *

WordPress Cookie Notice by Real Cookie Banner