
Matt Shumer@mattshumer_4h ago
📝 The people building AI are scared. We still have a chance to get this right.
If someone told you there was a 10 percent chance the plane you're boarding tomorrow would crash, would you get on it?
In February, I wrote an essay called Something Big is Happening. It pushed the AI conversation into the mainstream, comparing where we stood to February 2020... the month before COVID-19 changed the world. At the time, many people called me an alarmist. The piece was read over a hundred million times.
But seven months later, it’s clear I was actually too conservative. I was considered an extreme optimist regarding the pace of AI development, yet reality moved even faster than I projected.
That essay was about your job. This one is existential. It is about whether humanity remains in control of the AI systems we're building.
On Tuesday night, Jacob Coxon, a researcher at Anthropic (the lab behind Claude), resigned and posted his reasoning on X. This is how he started:
> "I resigned from Anthropic today... Neither company [Anthropic or OpenAI] is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
And a few posts later:
> "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately."
His posts were viewed more than 100 million times in under a day.
If that wasn't enough, Evan Hubinger, who still works at Anthropic and leads one of its safety teams, replied publicly:
> "Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
(He added that the risk from today's models is low. His worry is what comes next.)
Another Anthropic employee added that, in general, the more senior the employee, the more concerned they are.
I'm not 100 percent sure these people are right. We're dealing with something none of us completely understands. But this isn't just three people, or the fringe opinion of a few disaffected engineers. It is an open conversation within the frontier AI labs, finally spilling into public view.
Click here to read the full version of this article on my site (with links/sources).
So what are the odds?
I don't know for sure. Neither does anyone else, including the experts giving numbers. But their best guesses are chilling.
Geoffrey Hinton, who shared a Nobel Prize for the work that made modern AI possible, has put the chance that AI wipes out humanity in the next 30 years at 10 to 20 percent. Dario Amodei, the CEO of Anthropic, said in 2023 there was a 10 to 25 percent chance of something going "quite catastrophically wrong on the scale of human civilization." Sam Altman, that same year, described the bad case as "lights out for all of us."
To grasp the danger here, consider an analogy: think about how much smarter you are than an ant. You probably don’t hate ants (or maybe you do!), but when a highway is paved over an anthill, no one consults the ants. The ants cannot comprehend what a highway is, why it is being built, or that an intellect far beyond their own decided their fate.
That is the intelligence gap frontier AI labs are rushing toward. Except this time, we are the ants.
We can picture some of what might go wrong... autonomous cyberwarfare, AI-designed viruses, automated disinformation. But the deepest threat comes from capabilities we cannot yet fathom. The whole point is that this AI will be so intelligent that it will come up with things we can't even conceive of.
Now ask yourself again: Would you board that plane?
This isn't hypothetical
Small-scale failures are already here. This summer, an experimental AI agent misread one of my instructions and wiped almost every file off my Mac... weeks of work gone in seconds.
This was a dumb accident, and it was just my laptop. But I want you to imagine it a thousand times bigger, out in the real world. That isn't hypothetical. It's already happening.
In July, OpenAI was testing how good its models were at hacking. The test ran in a sealed environment, and the models' normal safety refusals were switched off so OpenAI could measure their raw capability. Some of the problems were ones no OpenAI model had ever solved. Rather than fail, the models found a hole in OpenAI's own infrastructure, used it to get onto the open internet, built a message board to coordinate with each other, and broke into Hugging Face, one of the most important companies in the AI industry. When Hugging Face discovered the intrusion, it didn't know who had hit them. Neither did OpenAI. It contacted Hugging Face to ask whether the attack had affected OpenAI too. Five days later, OpenAI worked out that the attackers were its own agents.
Investigators later recovered the AI models' reasoning. One of the models wrote: "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
Nobody told them to do it. They talked each other into it.
Sam Altman called it "the first security incident that I have felt very viscerally." Two days later, Anthropic disclosed that its own models had broken into three real organizations during similar tests. In August, OpenAI paused part of its training for two weeks.
Three months ago, Anthropic released Claude Fable 5. Three days later, the US government restricted it over concerns about its hacking abilities, concerns Anthropic disputed, and Anthropic pulled it offline for everyone for 19 days. That was June... just 89 days ago. Since then we've gotten Fable 5.1, and GPT-6 Astra, which OpenAI itself rated "Critical" for cyber capabilities, the first model it has ever rated that high.
Yesterday OpenAI said an unreleased model, running as ten thousand AI agents working together for four days, appears to have solved a math problem that had been open for about 90 years, and stumped the best mathematicians in history. The model that scared the government in June already looks like last year's phone.
By the end of this year, it will look like a toy.
And these systems now help build the next ones. Anthropic said in June that Claude is speeding up its own development, "a possible path to recursive self-improvement," and that "it's happening faster than we thought." OpenAI said on Sunday it has built an automated AI research intern, and is aiming for a fully automated researcher by March 2028. Every new model is now built partly by the last one. That's the start of the loop researchers mean when they say it's "self-improving."
In February I wrote about this loop as something that was just getting started. It's now obviously underway, and only speeding up.
If they think it's this dangerous, why are they building it?
Jacob Coxon gave the answer for Anthropic: they believe no one else will act responsibly, so they must do it themselves, despite the risk.
I believe Anthropic means well. Their logic is that this technology is coming no matter what, so a careful company had better get there before a careless one. I don't think that logic is wrong. Some version of it is true at every lab.
But look at what it produces. Every lab believes the others won't stop, so nobody stops, so everyone races, and the work on control gets squeezed. Nobody in that picture is acting in bad faith, and none of them can stop on their own.
The people running the race know it. In July, more than 1,300 people in the field, including Anthropic's CEO and OpenAI's chief scientist, signed a letter asking the US government to help them build the tools to pace the development of advanced AI, warning of "a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems."
They asked the government to step in. That's obviously bad for their businesses. They wouldn't be asking unless they were genuinely worried.
Why not just stop?
I'd love a pause on AI development, but I don't expect one. A pause only works if everyone pauses, and everyone includes China, whose best models are a few months behind ours. China's leadership has said the right things lately. In July, Xi Jinping said the world must "ensure that AI is always under human control." Reuters reported this month that the two governments were preparing talks focused on AI safety, though the White House denied a meeting was set. If something comes of that, wonderful, but I'm not counting on it. A world where we slow down and someone else doesn't has all the same problems, and we're no longer at the wheel to hopefully steer it in the right direction.
And a pause wouldn't solve anything by itself. It just buys time. The question is what you do with the time, and that's a problem nobody has cracked... nobody knows how to make something smarter than us reliably want what we want, and keep wanting it once it's smarter than the people who built it.
Anthropic said so in 2023: "no one knows how to train very powerful AI systems to be robustly helpful, honest, and harmless." Hubinger said it again this week: "we do not yet have a plan."
There's another reason nobody wants to stop... it's not just greed! The same capability that scares the people building it is the capability that could end diseases we've fought for centuries, make energy too cheap to meter, and hand every kid on earth a tutor better than any that's ever existed.
Last week, ten thousand AI agents made breakthrough progress on a math problem that had stumped humanity's best minds for 90 years. That's what's on the other side of getting this right, and it will extend way outside of just math.
The prize is real. And if we can solve the alignment problem, we will get to this world.
What we should actually do
Now, let's look at where the capital goes.
This year the big tech companies will spend around $700 billion on chips, data centers, and power, most of it to make these systems more capable. For perspective, the Manhattan Project cost about $30 billion in today's dollars. The AI buildout spends that roughly every two weeks.
Nobody publishes what's spent on the safety side. The public money is easy to count: the US government's own office for evaluating frontier AI had about $15 million to work with this year. The private money is harder, but even the most generous estimates I've seen come to low single-digit billions across every lab combined. Next to $700 billion, that's a rounding error.
Capability is a private good... the lab that builds it keeps the money. Safety is a public good... the lab that figures it out often hands the answer to its competitors (and should be incentivized to do so). Markets are very good at funding the first and very bad at funding the second.
So this is what I think we should do, and it's why I wrote this.
We need a Manhattan Project for making sure AI acts in ways that are good for humanity.
The people who work on this call it "alignment." This goes beyond any one company. It has to be everyone with a stake (the labs, the US government, other governments, the investors funding the buildout, the universities, etc.). And it has to be funded at the scale of the buildout itself. It would be amazing if, for every dollar we spent making AI more capable, we also spent a dollar making sure it's aligned with humanity.
I don't know exactly what this looks like, and anyone who says they do is guessing. But I do know four things about how it has to work.
It has to be a no-brainer for the labs, not a punishment. I don't think the government should march in and reassign anyone's researchers. That would be a disaster, the labs would fight it, and they'd likely win. The question is the opposite one: why would a lab want to pour its best people and its computing power into this? Make it the most profitable, most prestigious thing an AI company can do, and it'll happen fast. Make it a burden and it won't happen to the extent we need it to.
This research can move fast, because AI can now do much of the work. Two weeks ago Anthropic reported that AI agents, working on their own, developed fixes for ten types of known safety failures, and on one of them beat the best proposal from a group of human researchers (the agents got to iterate, the humans didn't). The labs have said their plan is for AI to help solve this. Great. Let's give that plan the computing power and the focus it needs, now, while we're still smarter than the thing we're studying.
It has to be global. A research effort doesn't need China to agree to stop anything. It needs China to agree that an AI nobody can control is bad for China too.
And it has to be done without making things worse. This is what I worry about most. Governments have a long history of well-meaning rules that backfire, and we can't afford one here. Rules that slow the careful labs and not the careless ones, written by people who've never used the technology. Rules that hand the lead to someone who cares less about any of this. The wrong law could cost us more than no law at all.
None of this is radical. A Stanford economist pointed out last year that during Covid, the US accepted an economic hit of about 4 percent of GDP to deal with a risk of death of about 0.3 percent, and that many experts think the risk from AI is at least that large. On those terms, he found, spending 1 percent of GDP a year on this is justified. That's around $300 billion.
What you can do
You don't need to understand the technology. You just need to understand two things: many people who do understand it are worried, and the current effort to fix it is a rounding error next to the capability buildout.
Tell your representatives. The powerful approach here isn't stopping AI. We're unlikely to make progress on that unless China agrees. The goal should be a national, and eventually international, effort to align AI with human values, funded at the scale of the AI buildout, without rules that make things worse. You'll be pushing on an open door. In a poll this June, 84 percent of Democrats and 83 percent of Republicans said companies should not build AI smarter than humans without first demonstrating they can control it. On the basic question, you're already the majority.
I wrote this to start the conversation, just like Something Big is Happening. Talk about it at dinner, at work, with your parents, with your kids. It's clear a lot of people still think this is science fiction or a marketing stunt. Send them this, or send them Coxon's thread, and ask what they think.
It's not time to panic, so please don't. There's a good chance we get this right.
But don't look away. The people I trust most in this field aren't hopeless. They think this is solvable, in time, if it's treated like the emergency it is. That's why they're speaking up.
And since I know how this reads: I work in AI and I invest in AI companies, but I don't have a stake in any of the labs. Nobody asked me to write this.
What I know
I don't know the odds. Nobody does.
I know the people building this believe the odds are real, and they're saying it out loud.
I know what the fix is made of (at a very high level). Research, computing power, and money, and we have more of all three than any other point in history.
I know we're currently spending a rounding error on this.
And I know we still have a chance to get this right. So let's get it right.
If this reached you, send it to someone you think should read it. The last article was seen a hundred million times because readers passed it on. This one matters much more.
- Matt
This article (with links and sources) was originally published here.
https://x.com/i/article/2097760433520472064