Tech News
← Home  ·  All topics

Grpo

1 GoKawiil brief on this topic

METR study: AI models underperform human researchers on novel RL post-training task

METR tested frontier AI models on a research task requiring them to invent a new post-training method that would outperform a strong GRPO baseline across question-answering and coding benchmarks. The human algorithmic innovation in the comparison outperformed the solutions generated by the AI models tested, according to METR's report.