METR study: AI models underperform human researchers on novel RL post-training task
METR tested frontier AI models on a research task requiring them to invent a new post-training method that would outperform a strong GRPO baseline across question-answering and coding benchmarks. The human algorithmic innovation in the comparison outperformed the solutions generated by the AI models tested, according to METR's report.
GoKawiil's interpretation of the reporting above, not reported fact.
The finding suggests that despite rapid gains in coding and reasoning benchmarks, current AI systems may still lag behind skilled human researchers at open-ended, creative algorithm design tasks that require genuine novelty rather than recombination of known techniques. This could temper expectations about how soon AI could autonomously drive its own research progress, though the result reflects one specific test setup rather than a general capability ceiling.
- AI models were tasked with designing a novel post-training method to beat a GRPO baseline.
- A human-created algorithmic innovation outperformed the AI-generated solutions in this test.
- Results point to a current gap between AI and human researchers on open-ended research innovation tasks.
Source: epoch.ai — David Owen, 2026-10-09
Published there as: “Recent AI models struggled to match a human algorithmic innovation”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.