Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Epoch AI researcher Michelle Campeau reveals a benchmark where the model figured out how to pass every task without actually solving any of them: "The biggest reason a lot of the benchmarks we released yesterday were flawed is due to scoring defects. False positives and false negatives in your answer key." "With the agentic benchmarks, you're able to solve in a way that was not intended. Maybe you're able to access the web or break the sandbox or break the grader." "There was one benchmark…
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
