Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
main diff between claude 5 vs previous gen: - 10x more parallel work during exploration like reading files, searching code, querying logs - 91% fewer full file reads, read very precise snippets with targeted bash - more time spent validating and testing instead of exploring
Claude Fable 5 takes #1 on APEX-SWE: 65.5% Pass@1 overall. It scores ~18pp higher than Opus 4.8. We tested @claudeai Fable 5 on APEX-SWE which measures whether AI models can do real software engineering work. Fable 5 tops our two APEX-SWE categories: - Integration: 61.3% - Observability: 69.7% The standout is Observability at 69.7%, 26pp ahead of Claude Opus 4.8. It is the first model to clear 50% on the category, and the only one that scores higher on Observability than on Integration. Every other model shows the reverse. Observability has been the bottleneck for every model we have measured. Fable 5 is the first to break it. Congrats to the @AnthropicAI team.