Compute-efficient pretraining and scaling to trillion-parameter models
(news.ycombinator.com)
1.
2.
>10x More Efficient Pretraining
(news.ycombinator.com)
3.
RSI Simulator
(news.ycombinator.com)
4.
NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute
(news.ycombinator.com)