Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
METR AI researcher @ReillyHaskins02 reveals how an AI agent can prompt-inject the model monitoring it by disguising its own messages as the user’s: "It's basically an open science question of how do you prevent LLMs from being susceptible to jailbreaks. I think basically we have no super good answer." "You can render the transcript in some more structured way, like a JSON, or having hashes corresponding with different roles, like the user or the agent. These can help prevent things like…
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors! https://t.co/7dSbXzV9QZ