Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
With low reasoning effort in our standard harness, GPT-5.6 Sol used reasoning tokens in 98% of the actions it took while playing ARC-AGI-3 public games. With the same low reasoning effort, GPT-6 Astra used reasoning tokens in only 27% of its actions (!). We saw a similar pattern with the semi-private games, suggesting this difference was not driven by public benchmark contamination and may instead reflect architectural changes in how the GPT models work.