Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 explains why one researcher misreading the rules pushed METR toward automatic safety enforcement: "There's obviously more human error involved if you let humans enforce things. If it's easy to do so, you should probably remove that vulnerability." "That's one thing that we're trying to do more in our org, is enforce policies automatically." "We have a central inference platform, Hawk, and every eval is running through that system, so we can implement someā¦
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ