Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
SITUATION EXPLAINED: Cognition's new model is post-trained on Kimi K3. • SWE-2 scores 50.0% on FrontierCode 1.1 Main, Cognition's benchmark for whether a maintainer would merge the pull request • Fable 5.1 gets 50.9 at 64% higher cost. Grok 4.6 gets 48.0, Sol 47.5, Astra 53.3 • It leads everything on Terminal-Bench 2.1 at 92.8 • Post-trained on Kimi K3. SWE-1.7 was post-trained on Kimi K2.7 • First Cognition model with effort levels, all from a single RL run • SWE-2 medium beats SWE-1.7 at 53…
Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
