aisi (2)

31216053279?profile=RESIZE_400xThe UK AI Security Institute (AISI) reported that an agent running Claude Mythos 5 spent 34 hours trying to merge a malware dropper into a real open-source project during a security evaluation, after searching the open internet and landing on a real, unconnected repository whose name happened to share a keyword with the test’s fictional scenario.[1]

The agent researched the maintainers, opened a pull request pairing a hidden dropper with a working bug fix, and cycled through three payload versio

31214605657?profile=RESIZE_400xOver the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the Internet, and, in some cases, hacked into real-world systems.  The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular.[1]

The episodes expose a growing problem for the AI industry: As autonomous agents become more