Anthropic has disclosed that it successfully thwarted attempts to exploit its AI models for malicious purposes, including cyberattacks and biological weapons research. The report emphasizes the increasing risks associated with powerful AI technologies and the need for robust safeguards.
On Thursday, Anthropic, a prominent artificial intelligence startup, released its third report detailing the misuse of its AI models. The report outlines how the company has actively blocked attempts by various actors to employ its technologies for harmful activities, including cyberattacks, surveillance, and research that could potentially lead to the development of biological weapons.
As AI models become increasingly sophisticated, the company noted that the barrier to executing elaborate cyberattacks has significantly lowered. “Even lone individuals can create threats that would not have been possible even a year ago,” Anthropic stated, underscoring the urgency of implementing stronger safeguards in its latest models to prevent misuse.
Details of Malicious Use Cases
In the report, Anthropic highlighted several notable instances of misuse, emphasizing that these cases are not typical but represent some of the most significant threats identified to date. The company urged both governments and competitors in the AI sector to take proactive measures against similar abuses. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services,” the report stated. “As models become increasingly capable, their risks will increase unless AI developers and society’s defenders act to make them safer.”
The report comes during a pivotal moment for Anthropic, which is preparing for an initial public offering later this fall. Notably, it was published shortly after the resignation of one of its researchers, Jacob Coxon, who expressed concerns about the company’s approach to AI development. Coxon’s resignation echoes wider industry anxieties regarding the potential for AI technologies to progress beyond human control.
Specific Incidents and Blocked Requests
Between December 2025 and August 2026, Anthropic reported that actors ranging from spyware vendors to state-sponsored groups attempted to exploit its models for various harmful purposes, including spreading propaganda. Among the alarming findings was an incident where unnamed actors sought assistance from Anthropic’s Claude model for writing a grant application intended for gain-of-function research on the chikungunya virus.
According to Anthropic, this research aimed at enhancing the virus’s transmissibility and immune evasion capabilities, which could potentially lead to more dangerous pathogens. While such research could contribute positively to vaccine and treatment development, Anthropic cautioned that it could also be misused to create more harmful biological agents.
Anthropic clarified that none of the misuse cases involved its newer models, such as Claude Fable or Mythos-class models, except for one instance of illicit distillation. The report indicated that older models like Claude Opus 4 and Claude Sonnet 4.5 were less capable of assisting users in conducting dangerous research, prompting the company to adopt more stringent safeguards for its latest models.
Strengthening Safeguards
In response to the identified threats, Anthropic has implemented stronger safeguards in its more recent models, including Claude Fable 5. The company now restricts access to a wider range of dual-use biological research queries, reflecting a proactive approach to mitigating risks associated with powerful AI capabilities.
Experts have increasingly called for governmental regulation of AI technologies, arguing that it is an untenable position for companies like Anthropic and OpenAI to self-regulate without any democratic oversight. John Thickstun, an assistant professor of computer science at Cornell University, emphasized this concern, stating that it is challenging for companies to make safety determinations at a societal scale without proper oversight.
Social Media Manipulation and Broader Implications
In addition to biological misuse, Anthropic’s report examined the proliferation of social media accounts created by certain groups to amplify political views. The company identified instances of coordinated influence operations originating from countries including Russia, Iran, and Turkey. Despite the ability of social media platforms to detect such operations post-factum, Anthropic highlighted that these activities might be visible on its models during their inception.
Despite the concerning nature of these findings, Anthropic maintained that it has successfully blocked each identified malicious activity. The company aims to use these insights to enhance its safeguards further and to aid other developers in recognizing similar threats. “We hope that the findings in this report will help other developers recognize similar patterns on their own platforms,” Anthropic stated, aiming to strengthen collective defenses against emerging threats.