Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Epoch AI researcher Michelle Campeau explains why flawed benchmarks aren't just a leaderboard problem, but a training problem that produces reward hacking: "If you task a model to do a task that is broken, it looks into other ways of completing the task in the way you prescribed, which might be reward hacking. If you wrote it poorly, that can lead to really bad and really scary outcomes." "With agentic benchmarks, the ways that things are failing aren't just scoring defects. It's…
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
