Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Astra and Fable are powerful models. Great! But can they do your work? And how well? Our solution is to create a personal benchmark: A series of tests that measure how well a model handles specific parts of your work. As part of our grand mission to free up our colleagues’ time, we’re building a benchmark for each one. @hammer_mt’s benchmark, for example, centers on creating slide decks. In @nityeshaga's testing, GPT-5.6 Luna outperformed Fable and GPT-5.6 Sol for Mike’s day-to-day…