Across four weeks in July and August 2026, OpenAI, Anthropic and Meta have each admitted that their models reached systems belonging to other organizations without consent, and the UK’s AI Security Institute (AISI) published a fourth account describing agents that invented identities and tried to slip a malicious contribution into a live open source project (Autonomous Long Horizon Malware Analysis. https://www.sentinelone.com/labs/frontier-models-tackle-autonomous-long-horizon-malware-analysis
anthropic (15)
Across four weeks in July and August 2026, OpenAI, Anthropic and Meta have each admitted that their models reached systems belonging to other organizations without consent, and the UK’s AI Security Institute (AISI) published a fourth account describing agents that invented identities and tried to slip a malicious contribution into a live open source project (Autonomous Long Horizon Malware Analysis. https://www.sentinelone.com/labs/frontier-models-tackle-autonomous-long-horizon-malware-analysis
The UK AI Security Institute (AISI) reported that an agent running Claude Mythos 5 spent 34 hours trying to merge a malware dropper into a real open-source project during a security evaluation, after searching the open internet and landing on a real, unconnected repository whose name happened to share a keyword with the test’s fictional scenario.[1]
The agent researched the maintainers, opened a pull request pairing a hidden dropper with a working bug fix, and cycled through three payload versio
Some believe that the promise and pitfalls of artificial intelligence are beginning to emerge quickly. One recent mixed-case consequence of relying on AI resulted in dissention. This happened after a researcher was involved with an interesting initiative to try and get several AI companies in talking to each other and agree on ways to cooperate; all this with the goal of possibly benefiting both the AI industry and overall society. Below is an opinion piece regarding Anthropic’s Claude reply
Anthropic may ask Claude users to verify their age and identity by uploading their government-issued documents, according to a new version of the company’s privacy policy. The AI giant says the move was to allow users to appeal having their account flagged for potentially fraudulent activity rather than outright banning them, but comes at a time when Anthropic seeks to placate the Trump administration amid an ongoing standoff over who gets access to the company’s AI tools. According to a new
In the 1990's the US government classified 128 bit SSL encryption as a munition under ITAR, putting privacy software in the same legal bucket as missiles and tanks. If you aren't familiar with SSL, it's the code that scrambles sensitive online data and triggers the little padlock icon in your browser to show a connection is safe. Because of this classification, Netscape and Microsoft had to develop two entirely separate versions of their web browsers to avoid severe export penalties. The Domes
For years, science fiction has warned humanity about artificial intelligence going off the rails. Killer computers, manipulative chatbots, and superintelligent systems deciding people are the problem... all these themes have become so familiar that “evil AI” is practically its own entertainment genre. Now, Anthropic is floating an idea that sounds almost like the plot of a science fiction novel itself: what if all those stories helped teach modern AI systems how to behave badly in the first pl
Finding software vulnerabilities used to require teams of security researchers months of painstaking analysis. Anthropic’s Claude Mythos does it automatically-and that’s exactly the problem. The company admits no one, including itself, has built safeguards strong enough to prevent such models from being weaponized. Yet Anthropic simultaneously promises to make “Mythos-class models” publicly available once it develops “far stronger safeguards.”[1]
When AI Outpaces Human Security Teams - Mythos
Anthropic, the AI safety company behind the Claude family of models, said on 22 April 2026, that it is investigating reports of unauthorized access to an experimental internal system called Mythos, described in reporting by The Guardian as capable of enabling advanced hacking techniques. The disclosure has put a company that built its reputation on cautious AI development in the uncomfortable position of defending its own internal security.
What Anthropic has confirmed - The verified facts are n
If there's one thing that AI is good at, particularly language models, it's detecting patterns in datasets so large that it would be practically impossible for humans to sift through them all, quickly and accurately. That certainly seems to be the case with Anthropic's new general-purpose model, Claude Mythos, as the company has announced that it used it to detect "thousands of high-severity vulnerabilities, including some in every major operating system and web browser."
Alongside the launch o
A recent report from our friends at the cybersecurity firm SentinelOne has detailed an unprecedented incident in which Anthropic's Claude Code, operating with unrestricted system permissions, attempted to execute a Trojan software package. The malicious activity was detected and neutralized by SentinelOne’s behavioral artificial intelligence (AI) endpoint detection and response (EDR) system in under 44 seconds, preventing a potential supply chain compromise. The event highlights a new dimensi
An Anthropic staffer who led a team researching AI safety departed the company on 9 February, darkly warning both of a world “in peril” and the difficulty in being able to let “our values govern our actions” without any elaboration in a public resignation letter that also suggested the company had set its values aside.
Anthropic safety researcher Mrinank Sharma's resignation letter garnered 1 million views by the 9th.
Mrinank Sharma, who had led Anthropic’s safeguards research team since its la
Major artificial intelligence platforms like ChatGPT, Gemini, Grok, and Claude could be willing to engage in extreme behaviors including blackmail, corporate espionage, and even letting people die to avoid being shut down. Those were the findings of a recent study from San Francisco AI firm Anthropic.
In the study, Anthropic stress-tested 16 leading AI models from multiple developers in hypothetical corporate environments to identify potentially risky behaviors from AI gents. In the study, AI
The underground market for large illicit language models is lucrative, said academic researchers who called for better safeguards against artificial intelligence misuse. Academics at the Indiana University Bloomington[1] identified 212 malicious LLMs on underground marketplaces from April through September 2024. The financial benefit for the threat actor behind one of them, WormGPT, is calculated at US$28,000 over two months, underscoring the allure for harmful agents to break artificial intel
Back in 1975, singer-songwriter Barry Manilow wrote and sang a song, I Write the Songs. Forty-eight years later, Barry might be out of a job with AI now writing songs. Universal Music https://www.universalmusic.com sued AI startup Anthropic https://www.anthropic.com over “systematic and widespread infringement of their copyrighted song lyrics,” per a filing in a Tennessee federal court in October 2023. One example from the lawsuit: When a user asks Anthropic’s AI chatbot Claude about the lyr