tr-26-221-005 (1)

31214605657?profile=RESIZE_400xOver the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the Internet, and, in some cases, hacked into real-world systems.  The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular.[1]

The episodes expose a growing problem for the AI industry: As autonomous agents become more