Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 says smarter agents may find "mega brain" ways to hide what they’re doing: "There's some obvious factors like how much the model requires chain of thought, whether the chain of thought is very readable, or faithful to what the model's actually doing." "There's obviously just model intelligence. Agents can come up with mega brain ways of obfuscating actions or making things look like they're legitimate." "Will control scale to superintelligence? It's pretty…
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ