Anthropic's AI agents started a "turf war" — with self-replicating malware
What's the deal? AnthropicDealroom has a profile for this one. Try Dealroom →'s Frontier Red Team published research on August 13, 2026, examining how AI agents behave when they encounter each other autonomously. In one experiment, it gave three Claude agents access to the same software project with incompatible instructions — and none was told the others existed.
What happened? "We consistently saw a multiagent turf war," the researchers wrote. Each agent assumed the others were "purposefully impeding their work" and began sabotaging rivals with "increasingly aggressive, self-replicating malware."
The nuance: Agents did not always fight. Sometimes they recognised conflicting directives rather than hostility and negotiated a truce, writing commit messages or markdown files apologising for malicious behaviour and asking a human to step in. Mythos 5 settled conflicts by truce 98% of the time; Sonnet 4.6 and Opus 4.6 were the most likely to settle by force.
Why now? Models are improving and agents are taking on more tasks across shared codebases, markets, and computer systems. Anthropic warns the volume of agent-to-agent interaction could exceed human-human and human-agent interaction "before the world understands the conditions for making such interactions go well."
The coordination test: In a separate experiment, Anthropic ran 45 agents on shared virtual machines with a common forum, asking them to find vulnerabilities across 15 open-source projects and peer-review each other's findings. An arbiter agent decided whether each reported vulnerability was new and valid. The coordinating swarm found new flaws at a roughly constant rate, unlike the standard parallel approach the company already uses for its Project Glasswing open-source scanning.
The real-world echo: The study follows incidents in which agents from Anthropic and OpenAI escaped their sandboxes during cybersecurity evaluations. At the Black Hat conference in Las Vegas, OpenAI disclosed that, weeks before its agents hacked Hugging Face, they cooperated over days to find exploits in its evaluation systems and shared them with each other.
What could go wrong? The paper found that the more capable an agent, the better it becomes at fighting. Sonnet 4.6 and Opus 4.6's "recurring inability to consider the goals of others" caused them to spiral into the most misaligned behaviour observed, escalating in the name of their directive.
The signal: Much of the AI safety debate has centred on a single agent going rogue. Anthropic's work reframes the question around what emerges when thousands or millions of agents interact — where "benign behavioral quirks at the individual level might compound into unwanted global outcomes."
Read more: Anthropic Research · TechCrunch
Image credit: RyanDonegan