Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
AVERI standards director @prpaskov reveals how they benchmarked Gemini without Google DeepMind ever seeing the benchmark or the benchmark owner ever seeing the model: "This is building on a pilot that was done at UK AISI with Anthropic two years ago, with a toy model and a toy benchmark. OpenMined provided the secure enclave in which both went. Neither side saw the other's." "What we did is we built on this, but using a Gemini model. We used a frontier model, and we also used a real benchmark…
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. → https://goo.gle/3St2xan
