Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Benchmarks have been lying to us about how good AI actually is. By 82%. A new paper just quantified something everyone in AI has felt but never measured: the model you are using is nowhere near the ceiling of what's actually achievable today. Here's the setup. Every benchmark you have ever seen reports one model, one run. But that's not how capability actually works: → Different models are good at different things → The same model gives different answers on different runs → Somewhere in…
