Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
you might think "oh this is just marketing"... but that's where you are wrong. spun up a fresh 8xb300 today and getting avg of 541 tokens/s and max of 1104 tokens/s when using it in claude code. at $54/hour, a bit expensive, but could absolutely be worth it for some tasks! https://x.com/casper_hansen_/status/2077…
🎉 SGLang v0.5.15 is out! We spent this cycle tuning GLM-5.2 NVFP4 for production serving, now hitting 500+ tok/s/user on 8x B300 and 450 on 4x GB300 (bs=1). We will put commands to run this at the thread below, and full technical details and instructions on a blog very soon 🫡 And we have some newly supported models: Hunyuan 3 (Hy3), Hierarchical Reasoning Model (HRM-Text), NVIDIA LocateAnything-3B, Baidu Unlimited-OCR, JoyEcho, and Qwen3.6. Here are highlights for this release: - Breakable CUDA Graph is now the default capture path - Native web search built in, powered by @ExaAILabs - Decode context parallelism for MLA models, including DeepSeek V3 - FlashInfer all-to-all for routed MoE - DeepSeek-V4: FlashMLA sparse prefill now on by default (>10% throughput on long context), plus a non-paged indexer for long-context prefill (>5% e2e) We welcomed 43 new contributors, and thanks again for our amazing partners and model makers: @NVIDIAAI @AMD @intel @Zai_org @TencentHunyuan @Alibaba_Qwen @deepseek_ai @Sapient_Int Now. MAX LOAD! MAX OUTPUT! 🚀
