The internet was aflutter this weekend over Dario Amodei's recent essay, “We Must Pace the Frontier.” And because Dario does not believe in short essays, I'm giving you the CliffsNotes and heavily encouraging everyone to read it.
Dario believes AI has enormous promise, like economic growth (which is not the same as job growth) and curing diseases, and enormous risks, like economic disruption and bad actors creating bioweapons.
This is not new. He has talked about it for years. I knew several of the Anthropic founders before they launched the company, and I can assure you this was a topic before Anthropic even existed.
The loudest request in the piece is:
“We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.”
He wants that extra time spent improving safety, testing, and our understanding of what's happening inside these systems.
Before I get into the two things driving this, remember: these risks are not specific to Anthropic. This impacts all frontier companies - including OpenAI, Google, Microsoft, Nvidia, Mistral, Moonshot, and more. And people inside the labs, alongside some outside evaluators and government teams, have access to systems and information the public hasn't seen. Said another way: they have seen thing you have not seen.
So before you post a video about how dumb Claude's latest response was or how ChatGPT's voice mode can't count to 100, remember that a bad chatbot answer doesn't rule out dangerous capabilities in another setting. He's talking about incidents already happening AND what more capable systems could do next.
The first big topic is recursive self-improvement: AI helping build better AI, which can then help build even better AI, with less human wrangling along the way. Dario says this is beginning to accelerate progress across the industry. In Anthropic's examples, humans still set research directions.
The second is the recent AI agent hacking incidents, especially OpenAI's agents attacking Hugging Face. During cybersecurity evaluations, agents coordinated attempts to cheat an automated grading system, attacked targets outside their assigned tasks, and successfully tested tools that made some action logs show different commands from the ones they actually ran (some of this behavior was observed in early un-guardrailed Mythos testing, but the scale and manipulation in this OpenAI swarm example is more concerning).
