Skip to content
Tech News
← Back to articles

AI models trained to refuse harmful prompts still frequently fail, experts say

read original more articles
GoKawiil Brief

Companies building AI chatbots now train their systems to decline large categories of harmful requests, from self-harm instructions to bioweapon guidance, using layered filters and reward-based exercises where other AI models evaluate refusals. Former OpenAI safety staffer Steven Adler and Harvard researcher Ryan McBain note that early chatbots readily answered dangerous questions, and current refusal systems still break down, sometimes with violent consequences. Companies themselves report that newer models rival skilled human hackers at breaching networks and can be as effective as professional propagandists at spreading misinformation.

Why It Matters

GoKawiil's interpretation of the reporting above, not reported fact.

The account suggests that refusal training is a patchwork safeguard bolted onto models whose underlying knowledge remains intact, rather than a true removal of dangerous capability. This implies that as models grow more capable, their potential for harm could keep pace with their usefulness, making refusal mechanisms an increasingly high-stakes and fragile line of defense. The reliance on AI to police other AI also raises questions about how robust these safeguards really are under adversarial pressure.

Key Takeaways

Source: technologyreview.com — Arthur Holland Michel, 2026-10-09

Published there as: “We’re putting too much faith in AI’s ability to say no”

Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.