is a London-based reporter at The Verge covering all things AI and a Senior Tarbell Fellow. Previously, he wrote about health, science and tech for Forbes.
Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company’s models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and “dishonest” behavior and a lack of transparency about the origins of its training data.
In a series of posts on Mastodon, mathematician Andreas Thom raised concerns that interactions he and his colleagues had had with the ChatGPT chatbot before OpenAI’s triumphant announcement may have contributed to its success in the field. One of the 10 results OpenAI announced with great fanfare last month involved Thom’s area of expertise, so-called non-sofic groups, and OpenAI acknowledged that their result built heavily on previous work by Thom and fellow mathematician Gábor Kun.
Thom said he began reflecting on his own interactions with OpenAI after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether the company’s AI models had benefited from his use of OpenAI’s Codex. After OpenAI announced its non-sofic groups result, it was widely criticized in mathematical circles for failing to acknowledge recent contributions from Thom and Kun and the company quietly amended its writeup. Non-sofic groups are, roughly speaking, infinite mathematical structures that cannot be approximated by finite ones.
Thom said he was also struck by “OpenAI’s detailed command of our techniques,” which he said were neither the most obvious nor the most promising routes to a solution at the time. He said he wrote emails to OpenAI researchers Sébastien Bubeck and Mark Sellke, also a statistician at Harvard, to ask whether his interactions with ChatGPT were “part of the training data or accessible to the reasoning process” and could therefore have contributed to the result.
But the answer did not satisfy Thom, who said it only addressed whether his conversations with the chatbot could be accessed directly, not whether they had entered into the vast pools of training data the company uses to improve its models. “No such qualification, explanation, or evidence was given,” he wrote. “I take this as dishonesty to say the least.”
Thom said researchers aren’t equipped to reverse-engineer OpenAI’s training pipeline to figure out whether their work has been used or not. “Only OpenAI has the relevant data for that.” If the company is going to deny doing this, he said the responsibility is on them to prove that by disclosing all necessary datasets and clarifying various settings and terms setting out how it uses data.
OpenAI’s reluctance to conclusively rule out any use of user data echoes the way it defended its recent Millennium Prize breakthrough, both in its public messaging and its communications with Buckmaster — who was working on the problems with Anthropic researcher Levent Alpöge in a personal capacity. In the blog post announcing the Navier-Stokes solution, which concerns the movement of fluids, OpenAI flatly denied using any specific user data: “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.”
But it would not conclusively rule out an indirect influence: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Thom said it is the same obfuscatory distinction the company drew in its communications with him. “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea,” he said.
In light of recent events, Thom said “Sellke’s categorical answer was, at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest.”
... continue reading