Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Google DeepMind researcher @iamtrask reveals how confidential compute lets AI labs take benchmark tests they can never see: "An AI lab develops a benchmark, evaluates their own model on that benchmark, hill climbs on it, and then publishes a statement saying their model is very aligned along some dimension." "This is an old problem in machine learning. People have been hill climbing on benchmarks since the dawn of benchmarks." "One fundamentally new paradigm is double blind evaluations. An…
FULL INTERVIEW: Andrew Trask says AI won't end with one giant model. He expects millions of models, ensembled and routed per prompt, to beat any single frontier model on quality and price, which would make AI look more like the PC and the internet than the mainframe. @iamtrask is a senior research scientist at @GoogleDeepMind and the founder of @openminedorg, which helped run the first double-blind evaluation of a frontier model. He joined @theojaffee to lay out who actually gets to audit the labs: 00:40 why embedded evaluators aren't enough 03:54 the "evil EAs" criticism of eval orgs 05:21 why labs only call people they already trust 08:22 a day in the life of an independent evaluator 10:10 how many evaluators belong inside the labs 12:09 will models just game the evals 13:42 the first double-blind eval on a frontier model 15:53 OpenAI's math release and sandboxes with no holes 20:40 why defining what's good could become a job 22:14 the case against one giant model winning 24:32 the proof from OpenRouter and Sakana