Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Neil on why AI spending today is different from the dot-com bubble: "A lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came. What's interesting about token consumption is that it's no longer speculative. People buy tokens because they're immediately valuable to them. You don't hoard tokens. You use them immediately. This is even different from what we had two years ago, where there was a supply crunch for…
Neil Movva (@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power. We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are. What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow. Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible. We discuss: - Latency versus throughput - Why there are no bad chips, only bad pricing - The end of kernel engineering - Buying chips and power no one else wants - New chip architectures - Nvidia lore + his contrarian view of the company - Open source and the frontier labs I learned a ton. Enjoy! TIMESTAMPS 0:00 Intro 0:38 Building a “Token Factory” 4:21 The Future of Background Agents 13:09 Nvidia and the GPU Stack 23:27 Chips, Memory, and Transformers 36:14 The Future of AI Training Data 44:32 Chip Scarcity and Compute Arbitrage 52:44 Reinventing the AI Data Center 59:01 Power and the “Scavenger Strategy” 1:10:10 Open vs. Closed AI