Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer should pay a premium for them." "But what I think is still underappreciated is, on the enterprise side, this is turning out to be existential as well." "We're in this world where it is still very unclear what ROI…
Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: https://www.youtube.com/watch?v=WO9c9qxDxzU @RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg