Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Epoch AI researcher Michelle Campeau on why so many AI benchmarks have errors in the first place: "A lot of people will write an answer and then write a question and send that pair to someone else to review. When you're horse blinders in on the answer, the answer can look right, but you don't understand what other types of answers would have also been right." "A lot of benchmarks these days are vibe coded. You can vibe code something that works to maybe 50% of what you actually want pretty…
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
