Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we’ve seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for the rest. They all share a general shape: Lossless Memory + Programmatic Analysis + Explicit Hypothesis Testing + Persistent State + Cheap Internal Compute. We’ve predicted (see v3 paper [2]) that as models get better…