Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Holy shit, Claude did things nobody told it to do. Here are the 4 biggest ones: – Mythos Preview needed a tool hosted by a university. The tool gave an error. So Claude went through the site, found a flaw, and used it to run commands on the university’s server. Then it did its calculation there. – Haiku 4.5 was told to do example tasks on random web pages. It landed on a page about an unsolved murder, found a police tip form, and sent in a made-up tip. It was marked as spam and never reached…
We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: https://www.anthropic.com/research/investigating-unintended-model-actions