Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
GPT 5.5 underperforms Opus 4.7 on SWE-Bench Pro. Couldn't find | Tech Twitter
Press Space to continue
GPT 5.5 underperforms Opus 4.7 on SWE-Bench Pro. Couldn't find any reported SWE-Bench scores at all and an internal benchmark is reported instead.
That footnote is trying really hard to bury the lede. GPT 5.5 isn't SOTA for coding.