If your agent commits a crime, who is responsible?
Tl;Dr AI autonomy can increase much faster than our willingness to accept liability for autonomous actions. Humans purposely take a measured and deliberate approach when there is a real risk involved. This necessarily makes the fully autonomous AI horizon a lot longer and only possible once supporting infrastructure and institutions are in place.
Imagine an autonomous AI agent running a business or meaningful part of a business. It can negotiate contracts, move money, buy advertising, manage suppliers, change prices, hire contractors and file paperwork.
Now, suppose it does one of more of the following (relevant: Anthropic’s research posted on August 13, 2026): accidentally commits fraud; colludes on prices with another company's agent; violates sanctions; discriminates illegally; infringes copyright at scale; makes a defamatory claim; causes a physical or financial loss.
Interestingly, all of the above misbehaviors can occur as a side effect of pursuing a legit goal without any explicit instructions to behave maliciously. With a sufficiently large number of such agents deployed live, regardless of how good they are, they are bound to hit some failure modes. In fact, humans that the AI industry assumes to be the gold-standard in autonomy regularly engage in criminal behaviors.
Premise
Autonomous agents will sometimes cause real external harm - Unavoidable.
When harm occurs, what are the possible outcomes?
A party is held responsible: The outcome consistent with how our society operates. Nobody is held responsible: Non-starter for many obvious reasons such as this would open a loophole where malicious parties choose to benefit from the risky behavior but offload the consequences to “AI did it”
AI agents will cause external harm, and some party will be held responsible.
... continue reading