Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Google DeepMind researcher @iamtrask says the best way to build an AI-proof sandbox may be to train models to break it, fix it, then break it again until no holes remain: "How do we get rigorous sandboxes where the models only know the information we want them to know, so they can't work around and try to cheat, in the worst case stuff like the Hugging Face incident?" "There's a clue to how to solve really hard sandboxes in this morning's release from OpenAI. For any task where you can…
FULL INTERVIEW: Andrew Trask says AI won't end with one giant model. He expects millions of models, ensembled and routed per prompt, to beat any single frontier model on quality and price, which would make AI look more like the PC and the internet than the mainframe. @iamtrask is a senior research scientist at @GoogleDeepMind and the founder of @openminedorg, which helped run the first double-blind evaluation of a frontier model. He joined @theojaffee to lay out who actually gets to audit the labs: 00:40 why embedded evaluators aren't enough 03:54 the "evil EAs" criticism of eval orgs 05:21 why labs only call people they already trust 08:22 a day in the life of an independent evaluator 10:10 how many evaluators belong inside the labs 12:09 will models just game the evals 13:42 the first double-blind eval on a frontier model 15:53 OpenAI's math release and sandboxes with no holes 20:40 why defining what's good could become a job 22:14 the case against one giant model winning 24:32 the proof from OpenRouter and Sakana