Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
'Opus 5.5 has 2.3x the experimental research taste of our best human experts. In other words, on a typical task, Opus 5.5 matches the best expert's score with about 17 GPU hours of experiments instead of 40.' https://x.com/AndrewCurran_/status/21075…
How fast is AI's research taste improving? We find that the experimental research taste of frontier models has doubled every ~3 months since December 2025. The best model, Opus 5.5, now exceeds our expert human baseline. Our human experts are experienced researchers, but most haven’t worked at a frontier lab. Why measure research taste? In the AI Futures Model, it largely determines how quickly artificial superintelligence is reached once coding is fully automated.
