Startup claims new pretraining recipe cuts compute 10x vs open-weight models
A research team reports developing a pretraining method that matches DeepSeek V4 Pro Base's performance using roughly 50 times less compute, at an estimated cost of about $0.5 million on GB200 chips. Scaling the same recipe up 10x further, to roughly $4 million, reportedly surpassed all publicly available open base models on perplexity benchmarks.