Researchers identify 'Matthew Effect' limiting RL training gains on hard math problems in LLMs
A new paper examines RL post-training of the Olmo 3 model on AIME math problems and finds that reported accuracy gains mask an uneven pattern: easy problems improve dramatically while the hardest problems, which the base model initially fails entirely, barely improve at all. The authors call this the 'Matthew Effect' and propose a technique called 'Never Give Up' to address it.