Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
I get asked all the time which models are best at querying data so believe me I was PSYCHED when I saw @Amplitude_HQ is starting to build their own benchmark for this exact question
We put GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 through our Amplitude Global Chat eval and mapped the outcomes on our quality/cost frontier. Our eval is 185 analytics tasks from simple to complex: building a chart, editing one, choosing the right event in a taxonomy, getting the correct date range, finding the property that most explains the variance in a chart. These evals are all based on real customer queries. We grade answers on correctness only. We measure cost and latency separately. Here’s where each new model landed (one run each so±2%): -Claude Opus 5.5: 83% correct at $0.70 per question. That's about 2 points behind our top score (GPT-6 Astra) at roughly a sixth of the cost. -GPT-6 Sol: 81% at $0.58. It scored a little lower than GPT-5.6 Sol, but was considerably cheaper (less than half the price). This is the cheapest model we've tested that clears 80%. -GPT-6 Luna: 76% at under $0.04, making it the cheapest model on the frontier. It's about 3x cheaper than GPT-5.6 Luna and achieves about the same score. The tradeoff is speed: its p95 latency was ~400 seconds, about 3x slower than Opus 5.5. Even with the inclusion of these new releases, GPT-6 Astra still gets the top score (84%) on our eval. The tradeoff is it costs $4 per question and was one of the slowest models we tested (900 seconds at p95).