Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 says AI safety’s window into model reasoning may be closing as models get better at acting without showing their chain of thought: "There's some evidence that points towards a downward trend in monitorability. Models can do things without chain of thought better than they used to be able to do, and it seems to be growing pretty fast." "Recent incidents have made people more aware of the need for monitoring, so probably more investments going into that…
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ