Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
TLDR; post-training can escape the pre-training ceiling if you have a dense reward system. a reward of either 1 or 0 is quite bad, and a reward that is continuous from 0 to 1 is more ideal. best quote: "When training these models utilizing a standard sparse reward, we indeed find that the base model’s probability of producing the desired behavior has a strong effect on the success of RL post-training"
