Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
The models saturated the last eval, so now we're stepping up the difficulty! How long until we're at 80%+? End of year?
Agents on Rails: Stage 2 is live. We wanted to find out: can you hand a model a real feature ticket and trust what comes back? The jump from atomic tasks to feature requests has interesting results… @OpenAI GPT-6 Astra is new to the leaderboard, and it came out on top: 35% of tasks solved, with 9-minute median runs, and relatively low cost, all at its default effort level: medium. @AnthropicAI Claude Fable 5.1 still performed well at second place, but came with a hefty price tag (almost 4x the cost of Astra). @GeminiApp 3.8 Flash was third place with a cost comparative to Astra, but took more 200 steps and longer at 27 minutes per run. At the bottom of the leaderboard, @OpenAI GPT-5.6 Luna, which did well in Stage 1 (46/63 tasks for $0.90), didn’t complete a single task in Stage 2 when the work required planning, migrations, testing, and completeness. Read the full benchmark report from @evilmartians here: https://t.co/yzMS6F0SWD