Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
The distinction between next token prediction and RL post-training is easily overstated. An RL pipeline can be reframed as a next token prediction objective and vice versa. What gets learned in either case is jointly determined by the learning objective and the training distribution. From the model's POV, it's gradient updates all the way down. Suppose you generate many successful reasoning trajectories and use this as data for a pre-training corpus. Next-token prediction then teaches theβ¦