Anthropic's Amodei Admits AI Firms Still Can't Explain How Models Think
A leaked resignation post from a junior Anthropic employee, who accused frontier AI companies of racing toward self-improving intelligence, spread widely and reignited public alarm over AI safety. A more senior engineer reportedly echoed internal estimates that the company's work carries roughly a 10 percent chance of catastrophic harm. In response, CEO Dario Amodei published an essay outlining a path toward safer AI, but conceded that researchers still understand only a fraction of how these models actually operate internally.