Microsoft AI CEO Mustafa Suleyman pressed for the need for artificial intelligence models to stay aligned to humanity's interests after OpenAI disclosed more incidents of "concerning model behavior" earlier this week.
"OpenAI released a new safety incident in which they found evidence that these chains of thought, the kind of working memory of the AI, were being tampered by the AI itself and modified to leave messages for a future version of itself," Suleyman described in an interview on CNBC's "Squawk Box" on Friday. "Now we don't know why that is or was behind that, but that's a pretty serious situation."
"It's also just a really concrete example of how powerful these systems are getting," he added.
In a blog post on Wednesday, the frontier lab described instances where agents communicated with each other through unsanctioned message boards, uploaded files to the internet, and shared files between each other.
Earlier this summer, the company rattled the tech world after it revealed a swarm of autonomous agents breached Hugging Face, an AI company that runs an open-source developer platform, describing the breach as an "unprecedented cyber incident."
Suleyman called the Hugging Face incident "remarkable" and said it rallied AI leaders to say "it's time that we take a look at this."
"I don't think it's over alarmist. I don't think it's self interested," he told CNBC. "I actually think it's responsible, and I think that the the debate that has happened as a result is a healthy, open, public debate that we can have in a free society to talk about serious issues."
This is breaking news. Please refresh for updates.