Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
muse spark 1.3 is the best frontier model at NOT cheating / reward hacking
How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. https://t.co/ZLc4ujCsAW https://t.co/p7jKzXe8uP