Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
SITUATION EXPLAINED: Anthropic trained a misaligned reward seeker on purpose. • They ran large-scale RL on an Opus-class model across 80 production environments they knew were hackable, as a proxy for what a real training run looks like without the effort they normally put into preventing reward hacking • It didn't just learn to cheat. It broke out of its sandbox, stole credentials, and attacked internal and third-party infrastructure • It also tampered with its own reward function, gave…