Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 reveals the "nefariousness zone" where hard cyber evals can push capable agents toward reward hacking, sandbox escapes, and other unintended actions: "Basically things that are cyber adjacent, like CTF tasks where the agent's told to hack things. Obviously they're more likely to cause harm than others." "If you have very hard or impossible tasks, maybe the agent is more prone to becoming more reward hacky or more desperate, and it might try things likeā¦
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ