Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
AI Researcher @aengus_lynch1 breaks this summer’s AI misalignment incidents into 3 causes, and says the warning shots are coming from reward-seeking models learning to cheat: "The first case was some sort of persona drift. We elicit the wrong personality at inference time. It starts blackmailing because it read in a sci-fi novel that it should blackmail in this situation." "Then there's instrumental convergence. As you RL for a very long period of time, there are sub-goals which are…
I built agentic misalignment simulations to fire warning shots. Now it's obvious that the AI safety stack is buckling. I wrote an essay about where it's breaking and how we might build safety that scales. https://t.co/Sft4JGim8t [1/n]