Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Some data points comparing GPT-6 Luna's performance on ARC-AGI-1 and 2 vs GPT-5.6 Luna: GPT-6 Luna used about 25% fewer reasoning tokens overall across public v1/v2 tasks and reasoning levels. Scores improved at some levels and declined at others. For example, across all v2 public tasks with max reasoning: - GPT-5.6 Luna scored 64.2%, averaging 152k reasoning tokens/task - GPT-6 Luna scored 61.8%, averaging 117k reasoning tokens/task On v1 public tasks at max, the score increased from 91.3%…
GPT-6 Luna from @OpenAI on ARC-AGI (Verified): - ARC-AGI-3: 0.19%, $241 (standard harness), 0.59%, $237 (provider adapter harness) - ARC-AGI-2: 59.3%, $0.062/task - ARC-AGI-1: 86.7%, $0.018/task GPT-6 Luna achieved similar scores to GPT-5.6 Luna at roughly 62% lower cost.




