Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc.
We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark.
As it turns out, the benchmark is perfectly…
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it…
This is the wrong approach. This kind absolute, overbearing take on AI regulation will prove to be actively counterproductive.