The risks emerging from frontier AI are real, and pretending otherwise isn’t a serious position. But neither is assuming that stopping development is the safest path forward. We need independent oversight capable of moving at the speed of AI, testing the systems that actually present meaningful risk while allowing the rest of the industry to keep building.
Sen. Bernie Sanders and Rep. Greg Casar announced legislation this week that would permanently ban the development and deployment of artificial superintelligence and temporarily pause advanced AI development until a new federal regulatory agency is established and has created rules for reviewing these systems.
I’ve spent some time thinking about their proposal because I agree with more of the underlying concern than I do with their solution.
The incidents we’re beginning to see involving frontier models and autonomous agents deserve serious attention. We’re giving AI systems more tools, more autonomy and more opportunities to interact with other systems. As those capabilities increase, there are legitimate questions about whether the companies developing the models can always predict how they’ll behave, particularly when agents operate over longer periods or encounter situations their developers didn’t anticipate.
I don’t think “trust the AI companies” is an adequate governance strategy for systems that reach that level of capability.
I also don’t think “stop developing AI” is an adequate answer.
My concern with the Sanders/Casar proposal isn’t that it takes AI risk too seriously. It’s that it makes a pause in development a central part of the answer, while putting much of the eventual authority into a new federal agency. I would rather see us build the technical capability to independently evaluate frontier systems while development continues, with the authority to intervene when the evidence from those evaluations justifies it.
That’s a different approach to governance, and some of my thinking about it comes from experience well outside today’s AI debate.
I’ve seen what governance looks like at scale
When I was a KCS Program Manager at Broadcom, I worked with a knowledge ecosystem containing more than 160,000 articles supporting a large global organization. Thousands of people had contributed to that environment over time, and one of the realities of operating something at that scale is that policies alone don’t protect you.
You can train people. You can establish publishing standards. You can tell everyone what belongs in the knowledge base and what doesn’t. At some point, though, you have enough content and enough contributors that you need ways to verify what’s actually happening in the system.
One of the projects we undertook was a large-scale audit for PII and sensitive information in the knowledge base. It wasn’t because we believed people were intentionally putting sensitive information at risk. The problem was scale. People make mistakes, standards change, old content remains accessible and information can end up somewhere it shouldn’t.
Stopping people from creating knowledge obviously wasn’t the answer. The knowledge base existed because the organization needed people contributing what they knew. The challenge was finding the risk, understanding it and improving the controls without destroying the value of the system in the process.
That’s stayed with me as I’ve moved deeper into AI.
Frontier AI is a vastly different problem, and the potential consequences are much larger. I’m not comparing a knowledge base to an advanced autonomous model. What carries over for me is the way I think about governance: if a technology creates enormous value but also introduces legitimate risk, the first question shouldn’t automatically be how to stop the technology. It should be whether we can understand the risk well enough to control it.
Sometimes the answer may be no. If that’s what the evidence shows, we need to be willing to act on it. But we need the evidence.
A pause sounds simpler than it actually is
The idea of pausing advanced AI development has an obvious appeal. If we’re worried that capabilities are advancing faster than our ability to understand them, slowing down gives everyone time to catch up.
The problem is that AI development isn’t happening inside one company or even one country.
The United States can pass a law. American frontier labs can be required to follow it. What happens if researchers or governments elsewhere continue moving forward? What happens if one group of companies agrees to slow development while competitors don’t?
I don’t think those questions are arguments for ignoring safety. They’re arguments for building a safety model that acknowledges the environment we’re actually operating in.
AI has economic, military, cybersecurity, scientific and geopolitical implications. Countries have strong incentives to develop it, and companies have strong incentives to compete. A regulatory strategy that depends on everyone voluntarily agreeing not to advance the technology strikes me as extremely difficult to sustain.
There’s another problem with stopping development that I think gets less attention: development is also how we’re learning where the risks are.
Each generation of models teaches researchers more about reasoning, tool use, autonomy, alignment, security and unexpected behavior. We find capabilities we didn’t anticipate and weaknesses that weren’t obvious before the model existed. That knowledge becomes part of how the next generation is tested and secured.
If we want to govern advanced AI intelligently, we need to understand it. That requires continued research, development and testing.
The challenge is figuring out how to do those things without blindly trusting whatever comes out the other side.
Not every AI system needs the same level of oversight
One thing I would strongly resist is building an AI regulatory structure that treats everything using a large language model as though it presents the same risk.
It doesn’t.
I’ve spent a lot of time working with systems such as ChatGPT, Claude, Grok and Gemini, along with AI-assisted development tools and other generative AI platforms. Businesses are using these technologies for knowledge management, customer support, coding, research, search, analytics and countless other applications.
Those systems can certainly create problems. Companies need governance around privacy, security, data handling, hallucinations, access controls and how employees use them.
But that’s a very different category of risk from a frontier system with advanced reasoning, substantial autonomy, persistent access to tools and the ability to take consequential actions across external systems.
The governance model should recognize that difference.
I don’t have a perfect formula for where the line belongs, and I don’t think anyone should pretend this is an easy threshold to define. Compute might be part of it, but I’m more interested in capability. What can the system actually do? How much autonomy can it exercise? What happens when it is given tools? Can it reliably circumvent restrictions? Can it conceal behavior from an evaluator? What happens when multiple agents interact?
Those are the kinds of capabilities that should trigger a higher level of scrutiny.
That approach also matters economically. We shouldn’t create a regulatory environment where a company using an LLM to improve customer support is treated like a frontier lab developing highly autonomous systems. Nor should infrastructure providers, data center operators and the broader AI ecosystem become collateral damage because our definition of “advanced AI” is too vague to distinguish where the meaningful risk actually exists.
Govern the capability that creates the risk.
Independent assurance makes more sense to me than a broad pause
What I’d rather see is an independent assurance framework specifically designed for frontier AI, potentially evolving into a worldwide, independent AI safety commission as international cooperation matures.
I don’t claim to have the organizational structure figured out, because independence is much easier to advocate for than it is to build.
Who funds the organization without gaining influence over it? Who appoints its leadership? How do the United States, China, Europe and other AI powers participate without any one of them controlling it? How does an independent evaluator attract people with enough technical ability to challenge systems being built by some of the best-paid AI researchers and engineers in the world?
Regulatory capture is a real risk as well. An oversight body isn’t independent simply because we put the word “independent” in its charter.
Those are difficult problems, but I think they’re the right problems to be solving.
Whatever the structure becomes, technical competence has to be foundational. The people evaluating frontier models need to include researchers, engineers, red-teamers, cybersecurity specialists, agent researchers and others who understand these systems deeply enough to challenge them.
They also need access.
Once a model crosses established capability and risk thresholds, independent evaluators should be able to test it before broad deployment. That evaluation shouldn’t simply reproduce the benchmarks the developer already ran. The point should be to find the circumstances where the safeguards don’t work.
Put the system into adversarial situations. Give it tools and increasing levels of autonomy. Evaluate it over longer periods. Test what happens when multiple agents interact. Look for attempts to circumvent restrictions, conceal behavior or accomplish a goal through a path the developers didn’t anticipate.
In other words, try to find the failure before someone discovers it in production.
If the evaluation identifies a problem, the response should depend on the problem. A developer might need additional safeguards, capability restrictions, monitoring, more testing or a staged deployment. A serious finding should require remediation and another evaluation before the affected capability is broadly released.
There should also be a point where the independent evaluator can say that a system isn’t ready.
I think that’s important because otherwise we’re just creating another advisory group whose recommendations companies can ignore when they’re inconvenient.
The distinction I would make from a broad development pause is that intervention should be tied to evidence about a system and its capabilities. If testing demonstrates that a model presents a serious risk we don’t know how to mitigate, holding that model back isn’t anti-innovation. It’s responsible engineering.
Fast engineering and external oversight can coexist
Aerospace provides an imperfect but useful example.
SpaceX is known for rapid iteration. The company builds, tests, learns from failures, makes changes and tests again. That development model has allowed it to move incredibly quickly compared with traditional aerospace programs.
It also operates within an external safety and licensing framework.
The FAA oversees commercial launch and reentry with responsibility for public safety. When Starship missions have experienced mishaps, investigations and corrective actions have followed, and the FAA can require safety issues to be addressed before return to flight.
That doesn’t mean the FAA designs SpaceX rockets, nor does it mean every failure stops SpaceX from developing the next vehicle. There is a boundary between the company doing the engineering and the outside authority evaluating specific risks that affect the public.
AI obviously isn’t aerospace, and I wouldn’t simply copy that regulatory structure. The failure modes are different, the development cycles are different and software can spread in ways a rocket can’t.
What interests me is the operating principle.
Rapid innovation and external oversight don’t have to be opposites.
There will be friction between them, and that’s probably healthy. Engineers will sometimes believe evaluators are being too cautious. Evaluators will sometimes believe developers are moving too quickly. The objective isn’t eliminating that tension. It’s making sure both sides are technically capable enough that the tension produces better decisions instead of bureaucracy.
For AI, speed becomes particularly important because an oversight organization that takes a year to evaluate a technology changing every few months isn’t much of an oversight organization.
It will always be behind.
That means AI governance has to innovate too.
Independent evaluators should be using AI to evaluate AI. Automated red-team agents, continuously updated adversarial tests, simulated environments and machine-assisted analysis could allow an assurance organization to test systems at a scale traditional regulatory review never could. Human experts would still be responsible for consequential judgments, but there is no reason the safety infrastructure surrounding AI should remain technologically static while the systems it’s evaluating improve exponentially.
We don’t have to start from zero
Pieces of this model are already beginning to appear.
The U.S. Center for AI Standards and Innovation has been working with frontier developers on model evaluation and AI security. In 2026, agreements with xAI, Google DeepMind and Microsoft expanded that work to include pre-deployment evaluations and targeted research on frontier AI capabilities and security.
I think that’s a promising direction because it demonstrates that development and independent evaluation can happen alongside each other.
It’s not the final model I have in mind. CAISI is part of the U.S. government, and its direction can change as administrations and policy priorities change. The agreements are also not the same thing as the independent international assurance structure I’m describing.
But they show that the underlying mechanism isn’t theoretical. Developers can build models while outside technical organizations evaluate them before deployment.
I’d rather expand our capability to do that well than make a broad pause our starting point.
Good governance should make AI easier to trust
There’s also a business reason to get this right that goes beyond avoiding catastrophic scenarios.
AI companies are asking enterprises, governments and individuals to trust increasingly capable systems with increasingly important work. Infrastructure companies are making enormous investments based on continued AI adoption. Businesses are redesigning workflows around AI. Employees are being asked to work alongside it.
Trust is becoming infrastructure.
A major failure involving an advanced autonomous system wouldn’t only affect the company that built it. It could produce a regulatory and public backlash across the entire AI ecosystem, including companies that had nothing to do with the failure.
That is why I don’t see good governance as something standing in the way of AI growth.
Done correctly, it protects the ability to keep growing.
Independent assurance gives developers another way to demonstrate that they’ve taken reasonable precautions. It gives enterprise customers better information when deciding what systems they’re willing to trust. It gives policymakers something more useful than company promises, and it gives the public evidence that somebody other than the developer has tried to find the weaknesses.
That’s valuable to an industry trying to move this quickly.
Internal safety teams still matter enormously. I’m not dismissing the researchers, engineers and security professionals inside AI companies doing this work today. They understand their systems better than anyone else, and independent evaluation would be much weaker without their participation.
But internal safety and external assurance solve different trust problems.
At a certain level of capability, I think we need both.
Where I disagree with Sanders and Casar
I understand the instinct behind the Ban Artificial Superintelligence Act. If lawmakers believe we’re approaching systems that humans may not be able to control, doing nothing would be irresponsible.
On that point, I agree with them.
Where I disagree is making a broad pause in advanced AI development the bridge between where we are today and whatever governance structure comes next.
I’d rather build the governance structure while we continue developing the technology.
Create meaningful capability thresholds. Require independent evaluation when frontier systems cross them. Give evaluators enough access to actually challenge the models. Require remediation when serious problems are found and give the oversight system enough authority to hold back a deployment when the evidence supports it.
At the same time, keep the rest of the AI ecosystem moving. Keep improving commercial models. Keep building infrastructure. Keep researching alignment and security. Keep learning from the systems we’re creating, because we’re going to need that knowledge to govern whatever comes next.
There may eventually be a capability so dangerous that development itself needs to be restricted. I’m not willing to rule that out simply because I’m excited about AI. Taking risk seriously means being willing to follow the evidence even when the conclusion is inconvenient.
I just don’t think a broad pause should be the default assumption before we’ve built the technical capacity to make those judgments well.
I’ve spent a good part of my career working at the intersection of knowledge, technology, governance and trust. What I’ve learned is that governance works best when it helps people use valuable systems responsibly. When governance becomes disconnected from how the technology actually works, people see it as bureaucracy and eventually find ways around it.
AI is moving too quickly, and becoming too important, for us to get that balance wrong.
I want AI development to continue. I want us building better models, better agents, better infrastructure and applications we haven’t imagined yet. I also want technically capable people outside the companies building the most powerful systems to have a meaningful opportunity to challenge those systems before we depend on them.
I don’t know exactly what the final international structure should look like, and I don’t think we need to pretend we do before we start building toward it.
What I am increasingly convinced of is that the choice isn’t between accelerating AI and taking AI safety seriously.
We have to learn how to do both.