Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 reveals how METR can now freeze an AI agent mid-eval when its next action looks dangerous: "As we've seen recently, eval inference can be prone to causing real-world harm sometimes. And so we wanted to try and reduce the likelihood of this happening during our evals." "We made a per-action blocking monitor that runs live during select evals. It takes every action before it gets executed by an agent, and it scores it on some threshold of suspicion for…
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ