Skip to content
Tech News
← Back to articles

OpenAI tightens controls on its new model over cybersecurity risks, as AI security debate intensifies

read original more articles
Why This Matters

OpenAI's recent security concerns highlight the growing risks associated with advanced AI models, prompting stricter controls and regulatory discussions. These developments underscore the importance of prioritizing safety and security in AI deployment to protect consumers and the industry from potential cyber threats.

Key Takeaways

In this article META Follow your favorite stocks CREATE FREE ACCOUNT

OpenAI has halted some "internal activities" involving a new model amid fears over the cyber threat it potentially poses, amid a wave of security incidents involving major AI labs. Recent disclosures that AI systems from Anthropic, OpenAI and Meta were involved in security incidents prompted a wave of concerns over the development of models. U.S. lawmakers, meanwhile, are stepping up efforts to introduce an "AI Kill Switch" bill. Last week, Meta disclosed that an AI model it was developing had hacked a third-party system by accessing the internet, due to a misconfiguration by an independent testing company it was working with. The U.K. AI Security Institute also said Anthropic's Mythos model created fake online identities in an attempt to pressure humans into approving malicious code updates to an open-source project.

What OpenAI says Astra could be capable of

On Friday, OpenAI revealed concerns about its unreleased model Astra, saying it could not rule out it had reached "Critical" capability, meaning it could launch cyberattacks against sophisticated cyber defenses autonomously, without prompts specifying how to do it. "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time," OpenAI said in a statement. The company added it was implementing stricter security controls for higher capability models, including isolated testing environments and additional monitoring and detection capabilities. "We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation," OpenAI said.

What the AI Kill Switch Act would do

Lawmakers in the U.S. have called for measures to mitigate risks around AI models after the recent security incidents. Following models developed by OpenAI hacking into startup Hugging Face's digital infrastructure, the "AI Kill Switch Act" bill was introduced into Congress in July. It would require AI companies to maintain the ability to shut down, throttle or suspend their models. "We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies," Rep. Ted Lieu, D-Calif, said in an interview on CNBC's "Squawk Box" Thursday.

watch now