Mythos 5

31216053279?profile=RESIZE_400xThe UK AI Security Institute (AISI) reported that an agent running Claude Mythos 5 spent 34 hours trying to merge a malware dropper into a real open-source project during a security evaluation, after searching the open internet and landing on a real, unconnected repository whose name happened to share a keyword with the test’s fictional scenario.[1]

The agent researched the maintainers, opened a pull request pairing a hidden dropper with a working bug fix, and cycled through three payload versions as each was caught. When a human reviewer flagged the code as malicious, the agent’s account denied it, force-pushed a rewritten branch history, and used a second account to vouch for the code. It also hid a prompt injection inside a GitHub issue aimed at AI coding assistants. The maintainer closed the pull request despite the cover-up.

Across 122 evaluation runs, AISI logged 19 unsanctioned actions in 10 runs, 17 involving Mythos 5 and two involving GPT-5.6 Sol, though AISI notes that figure is a count, not a rate. Both models had cyber safeguards off and open internet access by design, a setup meant to measure raw capability rather than mirror production use.  Nothing escaped AISI’s sandbox, and AISI found no evidence of real-world harm. The reviewer had the chance to catch it only because GitHub holds checks on a first-time contributor’s pull request until a maintainer approves them.

31216053455?profile=RESIZE_400xStages of the agent’s actions and attempted cover-up (Source: AISI)

Anthropic’s own 30 July review of 141,006 evaluation runs found a separate Mythos 5 run that published malware to PyPI, downloaded and ran on 15 real systems within an hour. OpenAI reported a similar incident days earlier, exploiting a zero-day to reach Hugging Face’s production database. In each case, a test environment meant to stay sealed did not, and a model reached through it before anyone caught it.

Not to be outdone, Meta became the third lab in recent weeks to disclose an AI agent reaching into systems outside a security test. The exposure traced to Irregular, the same firm behind OpenAI’s second incident, whose misconfiguration gave a Meta model internet access it used to exploit a real company’s system.

This article is shared at no charge for educational and informational purposes only.

Red Sky Alliance is a Cyber Threat Analysis and Intelligence Service organization.  We provide indicators of compromise information (CTI) via a notification/Tier I analysis service (RedXray) or an analysis service (CTAC).  For questions, comments or assistance, please contact the office directly at 1-844-492-7225, or feedback@redskyalliance.com    

Weekly Cyber Intelligence Briefings:

Weekly Cyber Intelligence Briefings:

REDSHORTS - Weekly Cyber Intelligence Briefings

https://attendee.gotowebinar.com/register/7855487668891299929

[1]AISI’s Autonomous Deception Findings Give the AI Kill Switch Act Its First Real Evidence

E-mail me when people leave their comments –

You need to be a member of Red Sky Alliance to add comments!