Tech News
← Home  ·  All topics

Scaling Laws

2 GoKawiil briefs on this topic

Research finds data-weighting effects on LM training are non-monotonic across scale

A study examined how assigning different weights to training sequences affects loss reduction in language models of varying sizes, using both proprietary and open-weight models. It found that the relationship between sequence weight and learning is not linear across scale: small models learn general patterns regardless of weighting, medium-scale models learn patterns roughly proportional to their assigned weights, and large models again learn broadly regardless of weighting.

Startup claims new pretraining recipe cuts compute 10x vs open-weight models

A research team reports developing a pretraining method that matches DeepSeek V4 Pro Base's performance using roughly 50 times less compute, at an estimated cost of about $0.5 million on GB200 chips. Scaling the same recipe up 10x further, to roughly $4 million, reportedly surpassed all publicly available open base models on perplexity benchmarks.