Tech News
← Home  ·  All topics

Chatbot Refusal

1 GoKawiil brief on this topic

AI models trained to refuse harmful prompts still frequently fail, experts say

Companies building AI chatbots now train their systems to decline large categories of harmful requests, from self-harm instructions to bioweapon guidance, using layered filters and reward-based exercises where other AI models evaluate refusals. Former OpenAI safety staffer Steven Adler and Harvard researcher Ryan McBain note that early chatbots readily answered dangerous questions, and current refusal systems still break down, sometimes with violent consequences. Companies themselves report that newer models rival skilled human hackers at breaching networks and can be as effective as professional propagandists at spreading misinformation.