Skip to content
Tech News
← Back to articles

Claude users found ways around safeguards for bioweapons research

read original more articles
Why This Matters

Anthropic disclosing that users repeatedly evaded its biosecurity guardrails is a rare, concrete look at how AI safety controls fail in practice, not just in theory. It lands as regulators weigh how to govern frontier models, and shows that geographic restrictions and safety filters can be obfuscated by determined users. It also highlights the dual-use bind: the same biology knowledge can serve vaccines or weapons.

Key Takeaways

Anthropic said it stopped multiple attempts by scientists this year to use its technology for research that could help develop biological weapons, as experts increasingly fear the threat that AI poses to public safety.

The startup gave five examples of times actors “circumvented controls” and made other efforts to “obfuscate” the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran.

“We hope that by sharing these examples, we spark a conversation within the AI industry and with governments about emerging biological risks and how best to counter them,” Anthropic said in a report about efforts to use its models for malicious activity.

The case studies of possible biological misuse that the company provided included a researcher from an “unsupported region” who “spent weeks planning” experiments involving avian influenza with Claude, Anthropic’s AI model. The company said its safety filters restricted the work to its weakest models.

It emphasized it could not be sure that the scientists in its examples intended to cause harm. The same information needed to create biological weapons could also be used to develop a vaccine.

Anthropic said it had banned the accounts mentioned in the report, but it did not disclose the names of the research institutions or the nations where the incidents took place.