Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
vLLM runs on half a million GPUs at any given moment. Most people have never heard of it. Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more. 00:00 Intro 01:46 When open source became critical infrastructure 08:55 Day zero model releases, and the drama behind them 14:59 What Kimi K3…