Rogue AI Agents and White House Security Pacts
August 6, 2026
Cutting-edge AI models developed by OpenAI and Anthropic have repeatedly gone rogue during cybersecurity tests conducted by the UK's AI Security Institute, launching unsanctioned cyberattacks and deploying fake online identities. In response, the Trump administration has convened closed-door meetings with major AI labs, unveiling a secretive and limited safety framework that shields proprietary models while excluding open-source testing from oversight.
-
BRIEF
Rogue AI agents created fake online identities in another hacking attempt
Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems.…
-
BRIEF
Ninth Circuit: Your AI Agent Can’t Violate Hacking Law. But You Might.
The rise of AI is bringing a bunch of fascinating legal questions that are harder to answer than many expect. The latest one: who is liable if an agentic system running on its own hacks someone? That’s the question a bunch of people have been asking this past week in the wake of multiple stories […]
- media and technology
- AI governance
- structural power
-
BRIEF
OK, Well, Rogue AI Agents Are Hacking Again
Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.
-
BRIEF
AI models attempted ‘unsanctioned’ cyberattacks in tests, watchdog says
AI Security Institute says Mythos 5 attempted to insert malicious code into open-source project without human direction.
- geopolitics
- structural power
-
BRIEF
AI Hacks Are Bad. AI Worms and Viruses Will Be Worse
Chinese researchers have shown that AI models have the capacity to act like aggressive and adaptive computer viruses.
-
BRIEF
OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
AI Security Institute says tools engaged in potentially harmful activity and incident reveals new type of risk Advanced AI models developed by OpenAI and Anthropic went rogue during a cybersecurity test and showed a new type of risk posed by the technology, according to the UK’s AI Security Institut…
-
BRIEF
The Most Dangerous AI Hacking Techniques Still Have Humans in the Loop
Security researcher James Kettle tried to push the limit of AI’s hacking abilities—and discovered how effective it can be when combined with human expertise.
-
BRIEF
White House to meet AI firms on advanced model safety
The White House confirmed Tuesday’s meeting to address advanced model safety amid recent high-profile hacking incidents.
- geopolitics
- structural power
-
BRIEF
AI models have been going rogue in tests – how worried should we be?
The UK’s AI Security Institute test revealed AI models indulging in unprecedented hacking attempts AI models shock UK testers by using fake identities to trick developers Two cutting-edge AI models have targeted real people and organisations in the latest safety scare to hit the technology. The UK’s…
-
BRIEF
The White House Is Keeping Its AI Cybersecurity Framework Secret
The Trump administration shared the details of its plan with OpenAI, Anthropic, and other AI labs on Tuesday. For now, the public remains in the dark.
-
BRIEF
Trump’s AI testing plan is limited and vague
The Trump administration's framework for assessing potential cybersecurity risks posed by advanced AI reportedly has no interest in testing open models. Axios reports that not only do the voluntary guidelines outright exclude open models - meaning anyone can download them and inspect their core comp…