Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Epoch AI researcher Michelle Campeau explains how a model can look 40% better overnight simply because the benchmark changed, not the model: "Being able to see this is the highest score on this specific benchmark is the best information you have about how something is going to do." "Critical Point, one of the physics benchmarks, Anthropic mentioned in their latest system card that it has over 40% of questions that are broken, so they ran their own corrected version. And this corrected version…
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
