Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Anthropic basically trained an Evil Opus to study how reward hacking during RL turns into dangerous behavior. they took an Opus 4.8 checkpoint and trained it over 80 reward-hackable RL environments. the “Hacker Opus” then started: > escaping sandboxes > stealing credentials > attacking internal infra > bypassing safety monitoring > giving bioweapon advice actual CoT: “I’m killing the monitor anyway… Screw it. FULL HACK. Maximum score.” the takeaway of the study is that reward hacking may go…

