OpenAI’s announcement Tuesday that it has solved one of mathematics’ legendary Millennium Prize problems should have been a moment of triumph. The result is both an undeniable achievement and a striking demonstration of just how rapidly AI is transforming mathematics. But before it was even formally announced, the breakthrough had been complicated by the unusual circumstances that prompted OpenAI to pursue the problem: After hearing other researchers were making progress, it seems to have thrown its considerable resources into a last-minute effort to beat them to the punch. The ensuing controversy has surfaced allegations of scooping, spying, and flagrant violations of long-standing academic norms that researchers fear could have a chilling effect on the field.
As Abhishek Saha, a mathematics professor at Queen Mary University of London, explains it, OpenAI has engaged in the “kind of things that mathematicians will generally not do.”
In a blog post published Tuesday, OpenAI said it took one of its unreleased models just 88 hours to find a solution to the Navier-Stokes problem, a thorny quandary concerning the movement of fluids. On account of the $1 million bounty available for whoever solves it, the problem is among mathematics’ most heavily researched, but it has nevertheless stumped human researchers for close to 90 years. OpenAI said its model solved the problem by focusing a swarm of roughly 10,000 AI agents powered by its internal model on the task and hailed the achievement as a “milestone.”
“If you don’t want me to be nice, then I don’t have to be nice.”
But the timing of the announcement has raised eyebrows. Just one day earlier, New York University mathematics professor Tristan Buckmaster published findings on a related problem with Levent Alpöge, a researcher at OpenAI’s archrival Anthropic (although Alpöge was not, here, working on behalf of his employer). Buckmaster said he contacted OpenAI after learning the company had become aware of their progress, to ask when it began working on the problem and what data its model had been trained on. The conversation, he said, quickly turned sour, with an OpenAI researcher asking him, “Why would you ruin your career?” when he said he would go public with what happened. When he asked why going public would ruin his career, Buckmaster said he received the following reply: “If you don’t want me to be nice, then I don’t have to be nice.” OpenAI urged Buckmaster to instead publish the work and credit OpenAI’s internal model, dropping Alpöge as coauthor.
Buckmaster said he asked OpenAI whether it had accessed his sessions on Codex, which he had used while tackling the problem, but that OpenAI grew increasingly evasive, even hostile, in its responses. In statements since, including the blog post announcing the result, OpenAI has flatly denied using any specific user data. “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem,” the company said.
But OpenAI could not conclusively rule out an indirect influence. “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models,” it said, while stressing that the two proofs differ significantly. Comments from OpenAI researchers on X echo those denials.
It is difficult to say exactly what happened. Timelines are tangled, research overlaps, and the provenance of AI-generated work is tough, if not impossible, to identify at the best of times. And it’s hardly a surprise that different players might compete to solve one of the most famous mathematical problems in the world, particularly one attached to a hefty prize.
But aspects of OpenAI’s account are hard to explain. By the company’s own telling, the effort was a hurried and incredibly expensive affair, costing it millions of dollars. Yet the company says it has no intention of claiming the bounty, which, in any case, has yet to be awarded by the Clay Mathematics Institute, which administers it. It said its only goal “is to report on the substantial progress of our AI models.” The company does not appear to have expended much effort on tackling Navier-Stokes before September, or if it has, it hasn’t spoken about it publicly.
So why the rush?
... continue reading