Research finds data-weighting effects on LM training are non-monotonic across scale
A study examined how assigning different weights to training sequences affects loss reduction in language models of varying sizes, using both proprietary and open-weight models. It found that the relationship between sequence weight and learning is not linear across scale: small models learn general patterns regardless of weighting, medium-scale models learn patterns roughly proportional to their assigned weights, and large models again learn broadly regardless of weighting.