Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
I like this benchmark. It feels more like how these models are used in day-to-day work. I’ll be interested to see how GPT-6 Sol fares.
So I tested Opus 5.5 (max) on real work. 2 repos. 105 planted bugs. Find and fix what you can. The results: GPT-6 Astra (max): 45 for $33.03 (n=3) Fable 5.1 (max): 43 for $77.55 Opus 5.5 (max): 43 for $60.49 Opus 5.0 (max): 27 for $51.33 Muse Spark 1.3 (max): 32.2 for $18.11 Anthropic is back in the game. More effort levels and attempts (n=3) dropping in this thread tonight 🧵