Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Mass General Brigham just published the largest clinical reasoning study of AI models ever. The FT headline says 80% misdiagnosis. The paper says 90%+ accuracy on final diagnosis. Both are true. The gap between those numbers is where all of medicine actually happens. The study, published today in JAMA Network Open, tested 21 frontier LLMs including GPT-5, Claude 4.5 Opus, Gemini 3.0, and Grok 4 on 29 standardized clinical cases. Each model was fed patient information in stages, the same way a…

