Anthropic the company detected and neutralized several misuses of its Claude models related to cyber espionage, the extraction of artificial intelligence capabilities, and the development of weapons-related technologies. These cases were detailed in the company’s threat intelligence report, which compiles incidents identified over the past eight months, the findings reflect a shift in how some malicious actors incorporate artificial intelligence into their operations, as the models can intervene at various stages while a human oversees the process.
According to Anthropic, this behavior goes beyond the traditional use of a chatbot; some operators have even resorted to multi-agent systems and frameworks capable of coordinating tasks and executing certain actions within broader cyber operations.
Anthropic detects a Russian-linked campaign against Ukraine
Furthermore, Anthropic identified an operation whose techniques matched those associated with Midnight Blizzard, a threat actor that the United States has previously linked to Russian foreign intelligence. According to the report, the group allegedly used artificial intelligence to support phishing campaigns, take over WhatsApp accounts and attack hotel Wi-Fi networks, the main targets belonged to government, military and diplomatic sectors of Ukraine.
AI was involved in virtually every phase of these operations, one of its most significant uses was the development of a mechanism to check when certain security tools detected malware, with this information, the system could modify the code and attempt to evade detection again. This case demonstrates how an AI model can be integrated into automated cycles that allow malicious tools to adapt during an attack.
Alibaba, DeepSeek and Moonshot appear in the Anthropic report
Meanwhile, Anthropic claimed to have blocked activities originating from seven China-based laboratories. Among the companies mentioned in the report are Alibaba, Moonshot, DeepSeek, and Xiaomi, attributed what it called its largest case of “illicit distillation” to operators linked to Alibaba. According to the US firm, the objective was to extract capabilities from Claude for use in training and improving Qwen models.
The company claimed to have identified over 151 million transactions attributed to Alibaba between May and July 2016. This activity reportedly reached nearly 3 million daily interactions through more than 3,500 accounts that Anthropic deemed fraudulent. Model distillation is a technique that allows smaller AI systems to be trained using responses produced by larger models. This process can reduce the computational resources and costs required to develop new systems.
The problem, according to Anthropic the identified operations sought to access Claude’s capabilities through mechanisms that violated the established conditions for using its models.
DeepSeek and Moonshot allegedly used customer conversations
Anthropic also accused Moonshot and DeepSeek of using another method to obtain information generated by Claude. According to the company, both firms channeled real customer conversations through their models. Some of these conversations contained sensitive information, and the responses produced by Claude were subsequently incorporated as training data.
The procedure would be different from a campaign based solely on millions of automated queries; in this case, real user interactions would have served to generate responses from Claude that could then be used to train or improve other models.
China’s Foreign Ministry told Reuters it was unaware of the Anthropic report, adding that the Chinese government supports the development of artificial intelligence for positive purposes and rejects both the distortion of facts and accusations directed against the country.
Claude was also used in weapons-related activities
In addition to cyber threats, Anthropic identified uses of Claude related to the development of software for conventional weapons and other associated systems. The company documented activities linked to firearms, missiles, armed drones, ammunition, and target selection and control systems. Some operators also reportedly used Claude to gather intelligence or support procurement processes related to weapons programs in China, Russia, and Yemen.
The report also addresses potential risks in the area of biological weapons. Anthropic cited five cases of scientists who used its models in ways that, according to the company, could contribute to this type of research. In one case, a researcher allegedly used virtual private server infrastructure to access Claude from a region where the service was unavailable. For several weeks, the researcher used the model to plan experiments related to the adaptation of avian influenza to mammals.
AI gains autonomy in cyberattacks
Finally, the cases described by Anthropic demonstrate an evolution in AI-related threats. Models are no longer limited to simply answering queries, but can participate in much broader task chains. This capability allows for the automation of certain stages of reconnaissance, code development, malware adaptation, and other processes involved in a cyber operation.
Anthropic also reported that it detected and neutralized activities related to affiliates of ShinyHunters, a cybercrime collective associated with attacks against large organizations. Jacob Klein, head of threat intelligence at Anthropic, explained that the capabilities of these models have increased over the past year. This advancement expands their legitimate applications, but at the same time forces developers to identify and contain new forms of abuse.
Source: Reuters
Photo: Shutterstock