When Daniel Evanko asked a scientist whether they had used artificial intelligence to write their peer-review report, he didn’t expect a confession. Most researchers don’t reveal AI help, says Evanko, who is the director of journal operations at the American Association for Cancer Research (AACR).
How to manage AI risks while reaping the benefits
But this time was different. “Wow, you guys are good!” the reviewer wrote back, admitting that he had used a large language model (LLM) after running out of time.
Evanko had a secret weapon. He and the AACR deploy a commercial AI-detection tool called Pangram, because of concerns over the number of peer-review reports submitted to their journals that seem to use AI without disclosing it, contrary to the publisher’s policy.
After years of disappointing results, multiple firms now claim that software can reliably distinguish between AI-written and human-written text. One is Pangram Labs, the New York City-based start-up that makes Pangram. “Detect AI-generated content with 99.98% accuracy,” the firm says on its website.
Scientists and research organizations are among those using the tool to spot AI’s traces. One in eight biomedical articles last year contained some AI-generated text according to Pangram, a study reported in January1. In June, the premier computer-science conference NeurIPS announced that it rejected 18% of submissions after screening them with Pangram. Users of the preprint server arXiv can now check Pangram’s verdict on any article there, by visiting a mirror site called alphaXiv that has installed the tool. And the University of Chicago in Illinois says that it has started using it to vet students’ coursework.
Pangram’s co-founder, Max Spero, says he wants to help everyone spot when text is AI-written. “If it’s taboo to call out that somebody’s using AI to write, then I think we’re going to see a lot more people shirking their jobs and letting AI replace themselves. We’re in a really critical time of setting norms,” he says. Spero has personally called out journalists whom Pangram suggests are using AI and, on one occasion, even flagged the Pope’s social-media posts as AI-written. This July, Pangram was integrated across the popular blogging platform Substack, allowing readers to see whether it deems posts to be AI-written.
Universities are relying on AI-detection software to catch cheating. How well do the programs work?
Pangram isn’t the only firm reporting remarkable results. GPTZero, a competitor also in New York City, says that it has 99% accuracy and provides “the most precise, reliable AI detection results on the market”. Five computer-science conferences and three universities have signed up to use it so far, says the firm’s chief technical officer, Alex Cui, and others are piloting it.
These AI detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, independent analysts say, although no tool can be perfect. And they are “good for screening out places that are pumping out slop”, says Tim Requarth, who studies science communication at New York University’s Langone Health centre in New York City.
... continue reading