In the last 2 months we launched 8 new models, 5 of which hit the #1 spot.
We now have the flywheel spinning. For each of image, voice and transcription we've been releasing a model every 7 weeks on average. That patient and relentless hill climbing is paying dividends.
All our focus has been on optimizing both quality and cost, something I wrote about a few months back: the frontier performance curve.
MAI-Transcribe-2 is the fastest, most accurate, and cheapest transcription model in the world.
MAI-Image 2.5 / 2.6 / 2.6 Flash all debuted at #1 on the Artificial Analysis leaderbaord and achieved the best quality-cost optimization out there.
MAI-Cyber in MDASH landed at #1 on CyberGym, while saving 50% of the costs versus other frontier models.
They’re also having a good impact in our products. MAI-Code-1.1-Flash launched in GitHub Copilot 4 weeks ago, and it now accounts for a third of the traffic for small models on GHCP and growing rapidly. In Excel, it's already better than Luna on our benchmarks.
Return on Tokens is now the name of the game, which is why we've had a relentless focus on nailing both quality AND performance. Getting this balance right is what our enterprise customers really care about as they use more tokens.
It’s an incredible time to be building. We know we still have a way to go, but its been great to see the velocity accelerating.
