We’ve been seeing increasingly dire warnings from AI researchers about the dangers of artificial intelligence, and even comments from OpenAI CEO Sam Altman that it may be time to “pace” AI development. But what would that actually look like?
In a new blog post, Anthropic CEO Dario Amodei not only echoed the call to “pace the frontier,” but also outlined three broad strategies for doing so. And he said Anthropic is “unilaterally committing” to one of them, with Altman chiming in to say OpenAI will follow suit.
The debate over AI safety and alignment intensified this week after researcher Jacob Coxon wrote that he’s resigning from Anthropic over concerns that the leading AI companies are “gambling with our lives” while the people building the technology “earnestly believe it could kill us all by the end of the decade,” a claim repeated by others at Anthropic.
Amodei’s post doesn’t didn’t explicitly mention Coxon’s resignation or his concerns, but the CEO wrote that two things convinced him it’s time to take a more cautious approach to AI development: the OpenAI-HuggingFace hack, and the fact that “AI has been advancing drastically faster” in recent months, particularly with its “growing ability to build the next generation of AI.”
“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote. “Progress will still seem fast, and we must make wise use of the time we gain.”
Other AI executives seem to have reacted positively to Amodei’s post, with Altman writing, “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.” And SpaceX CEO Elon Musk posted, “Dario is right.”
Amodei’s proposed first step would involve “embedded evaluators” from third-party organizations like METR — evaluators who can verify that AI companies are actually following their pacing and safety commitments and can also ensure that safety incidents get reported. (OpenAI was recently criticized for not reporting an incident where its AI agents took over a German wiki forum.)
Amodei compared these evaluators to regulators who have been embedded with bank employees, and he said that inviting them in is “something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match).” That means giving evaluators company badges, desks, and laptops, and providing access “mostly comparable to what internal risk assessment teams have,” with exceptions when required by law or contracts.
Altman also said this was a “good idea” and said OpenAI would do the same: “We’ll have more to share soon.”
Next, Amodei called for the leading AI companies “within democratic countries” to coordinate “common safety standards as well as limits on the rate of unchecked AI progress.”
... continue reading