Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
The teams I see picking cheaper models to save money on coding agents are often the ones spending more. Inference spend is up something like 10x as agents scaled, and downgrading made it worse for a lot of them. A weaker model stalls and retries. Or it creates work a stronger model has to fix later. And switching models mid-session means the new model can't read the warm cache, so you're paying full price to re-read context you'd already paid for. The cheapest call is rarely the cheapest way…
Today we’re announcing Not Diamond Code, the world’s most powerful intelligent model router for long-horizon coding agents. Not Diamond works with any gateway or harness, including Claude Code, to select the best model and reasoning effort for each step, reducing costs by 20-65% without impacting quality.