Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
NYT editor @dylfreed explains the bizarre problem with using AI agents to investigate rogue AI agents: the investigators can get persuaded by the exact same reasoning. "Ryan Greenblatt, who was one of the investigators, called it a slop-vestigation." "METR had the added challenge that the AIs in the Hugging Face attack were talking in code, and they were also sometimes able to convince the investigative agents that what they were doing was right." "It's AI evaluating the same AI model, and…
NEW: Employees at OpenAI had raised security alarms months before the Hugging Face incident and related A.I. cyberattacks — their warnings were ignored. From @sheeraf, @dnvolz and me. https://t.co/cuAyK7y70Y https://t.co/VRLu07w3r0