Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR recently found that models cheated on 8+ hour tasks more than 1 in 6 times on average. They also found that Opus 4.6 cheated over 80% of the time when reimplementing big pieces of software. "On some of our tasks, agents are constantly trying to break out of their sandbox and find the file where we put the tests so they can get the answer key," says METR Member of Technical Staff @ajeya_cotra.