Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Inherent researcher @DamonFalck on why the industry is wrong that GRPO needs verifiable tasks: they made it work on paper replication with a hindsight based rubric judge that scores individual turns "Everyone's trying to do GRPO RL for longer and longer horizon tasks, and they're running into problems. The solution has been, well, we need to constrain to verifiable tasks." "We decided that really wasn't the route we're taking, so we've had to come up with a bunch of innovations to make this…
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵 https://t.co/nA6ylMNrvj