Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Earlier this week, we showed that GPT-5.6 and Claude 5 exhibit very different long-context TTFT scaling, which corresponds to their different pricing structures at long context lengths. OpenAI recently released GPT-6 Astra, which shows the same pattern of increased API pricing beyond 272k input tokens. We collected additional latency measurements, which show a similar curvature to that of GPT-5.6 models.
OpenAI GPT-5.6 models and Anthropic Claude 5 models have different pricing structures at long context lengths. GPT model costs increase in price past 272k input tokens, while Claude model costs remain fixed. Does this reflect an underlying difference in the architecture of these models? Our measurements of serving latency suggest so. We studied time to first token (TTFT) on these models and how it scales with increasing context length. We found a significant difference in how they scale, with GPT showing a noticeable quadratic component, while Claude models remain closer to linear.

