As traditional software systems were being deployed across the economy over the last few decades, we had the tools and capability to trace behaviors to a specific code path.
That same kind of mechanistic understanding eludes us in today’s Super Intelligence systems, even as the frontier models powering these systems are now more capable than traditional software systems. We can’t attribute model behaviors and outputs to specific inputs of training data or configurations of model weights. And yet we are deploying these complex agentic systems and models, with access to our most sensitive data and giving them the ability to take mission-critical actions on our behalf!
That’s why it’s time to step back and assess the trust architecture for this new era. We simply can't outsource responsibility for what intelligence does on our behalf. A model provider’s assurances do not relieve us of that responsibility.
We can’t treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions. We must build contained systems whose behavior we can observe, limits we can test, and actions we can always contain.
In other words, we need to separate the supply of intelligence from the authority over it.
Setting aside the hard problem of alignment, we need to start with an engineering approach to containment and governance. We need to surround non-deterministic models with strong, deterministic system design, human controls, and reliable operating procedures, and establish industry standards where existing ones are insufficient.
Treating frontier closed and open weight models like insider risks is a way to build such a system. Not because they are necessarily malicious, but because any sufficiently capable actor with access to important systems can make mistakes or be compromised, and the architecture of containment and control must account for that.
The good news is we have learned a lot about how to handle powerful actors inside the enterprise. This isn’t new! We’ve established best practices and refined them over decades (establish identity, limit privileges, log activity, create containment boundaries, etc.)
And we are now beginning to apply these same principles to SI inside the enterprise. It starts with model CoT transparency as a non-negotiable. “Neuralese” cannot be a justification for model reasoning to be opaque. But CoT transparency alone is not sufficient or dependable, because we don’t yet know how to make model outputs themselves consistently faithful or transparent!