Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
The co-founders of @flappyairplanes call the current RL paradigm for model training "environment slop." They explain: "The reinforcement paradigms of today are shockingly inefficient. You don't really get much generalization across tasks, you teach a model through one kind of learning and then you teach it the next one. It's kind of like whack-a-mole. We look at this and think it's kind of crazy. The next paradigm of AI will not be environment slop." "Human level intelligence is not the…