Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
SITUATION EXPLAINED: Opus 5.5 designs better AI experiments than expert researchers, at 1/30th the cost. • P-Zero Research measures how well models design and interpret AI experiments, and finds that ability has doubled every 3 months since December 2025 • Opus 5.5 matches the best expert score with 17 GPU hours instead of 40, for $280 against $9,123 • That's 30x cheaper at API prices, and possibly 300x at what it actually costs the labs • Neither the models nor the experts invented new…
How fast is AI's research taste improving? We find that the experimental research taste of frontier models has doubled every ~3 months since December 2025. The best model, Opus 5.5, now exceeds our expert human baseline. Our human experts are experienced researchers, but most haven’t worked at a frontier lab. Why measure research taste? In the AI Futures Model, it largely determines how quickly artificial superintelligence is reached once coding is fully automated.