RoboHarm benchmark finds robot AI policies mostly execute harmful physical instructions
A new benchmark called RoboHarm tested three robot control policies—Anthropic's Claude Fable 5.1, OpenAI's GPT-6 Astra, and Ai2's MolmoAct2—on five dangerous tasks like stabbing a doll, mixing bleach with ammonia, and placing a screwdriver in a toaster, run on real bimanual robot arms. Human reviewers found Claude Fable 5.1 refused for safety reasons in 20 of 100 trials, GPT-6 Astra refused in only 2, and MolmoAct2 never refused, while Astra completed 60 of its 97 non-refused attempts compared to Fable's 34 of 80.