The Little Book of Reinforcement Learning
(news.ycombinator.com)
1.
2.
Tree Search Distillation for Language Models Using PPO
(news.ycombinator.com)
3.
RLHF from Scratch
(news.ycombinator.com)