Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Interesting results here. This is why I expect more agent workloads to run on blended models. Pareto 26.9 from @TheUnbiasedCo sends requests to several frontier and open models and keeps the best answer. In the new eval of 30 agent tasks, Pareto tied GPT-6 Astra for first place at about 1/3 the cost per successful task. It also finished tasks faster than DeepSeek V4 Pro and GLM 5.3 Flash.
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵 https://t.co/vSlEYGJbfb