Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Reasoning From Scratch: Reinforcement Learning with Verifiable Rewards (RLVR) round 2. Covering clipped policy ratios, KL loss term, format rewards, and other GRPO tips & tricks. 00:00 Introduction and recap 01:52 Interpreting basic GRPO training metrics 06:34 Planned improvements to GRPO 08:58 Running longer training jobs with Python scripts 13:39 Running the baseline GRPO training script 17:29 Loading and plotting training logs 19:29 Diagnosing unstable training 23:55 Evaluating checkpoints…