AI Security Breach: Anthropic Blocks Global Cyber Threats
Artificial intelligence developer Anthropic has released a concerning new threat report detailing how its Claude AI models were targeted by…
Featured image for news article
Artificial intelligence developer Anthropic has released a concerning new threat report detailing how its Claude AI models were targeted by state-aligned actors for malicious purposes. The findings reveal a sophisticated effort to exploit generative AI for weapons development, cyber-espionage, and large-scale psychological operations across the globe.
In one of the most alarming discoveries, Anthropic reported an attempt to utilize Claude for the creation of missile guidance software. Operators in northern Yemen allegedly tried to replace human engineers with the AI, assigning specific roles to different model instances to draft complex flight-control algorithms for guided rockets and long-range ballistic missiles.
Although the company’s internal safeguards successfully blocked most of these requests, some instructions managed to bypass filters. The actors attempted to evade detection by fragmenting tasks across multiple sessions, ensuring that no single prompt triggered a security alert. While Anthropic noted an unsuccessful test-fire, there is no evidence that a functional weapon was successfully produced.
State-Linked Espionage Campaigns
Beyond weapons manufacturing, the report highlights the use of Claude in sophisticated cyber-espionage. Anthropic identified a campaign linked to the Russian actor Midnight Blizzard—also known as APT29—which reportedly leveraged automated AI workflows to conduct phishing and data theft. These operations specifically targeted Ukrainian and European diplomatic entities, including various drone manufacturing firms.
Simultaneously, the company disrupted an offensive program originating from university students in Hunan, China. This group allegedly used Claude as an orchestration layer to probe government and corporate networks throughout the Middle East, Europe, and Southeast Asia. Anthropic responded by banning the associated accounts and implementing enhanced monitoring protocols.
Psychological Warfare and Surveillance
The report also sheds light on Iranian state-aligned activity, where three accounts were found using Claude for covert influence and psychological operations. These efforts were linked to entities such as the Islamic Culture and Communications Organisation and a seminary command room in Mashhad, both of which disseminated content aligning with the Islamic Revolutionary Guards Corps’ (IRGC) narratives.
“The company said it has no evidence the group managed to field a working weapon, though Anthropic claimed the operators appeared to have conducted an unsuccessful test-fire.”
Furthermore, Anthropic identified industrial-scale operations that used the model to build structured target profiles. By mapping individuals based on location, political leanings, and demographic data, these actors sought to enhance the precision of their influence campaigns. Anthropic described this as the most operationally mature case encountered during the investigation period.
To mitigate these risks, Anthropic has taken decisive action by banning the compromised accounts and sharing vital threat intelligence with public- and private-sector partners. These incidents underscore the dual-use nature of generative AI, highlighting the ongoing tension between technological innovation and the necessity of robust AI safety and security measures.
