Anthropic is trying to move oversight of frontier AI out of the realm of voluntary pledges and into binding law, and that shift now sits at the center of the policy debate in Washington and Sacramento. The company is urging Congress to make independent safety testing mandatory for the most capable models before deployment, including powerful open weight systems whose parameters can be copied and redistributed, and is tying that demand to a broader framework for managing catastrophic risk and economic disruption.
How we got to mandatory testing
For most of the last decade, AI safety has been governed largely by company policies, academic norms and scattered government guidance rather than enforceable rules. As models became more capable and started to show potential for serious misuse in areas such as cybersecurity and biological threat modeling, regulators and legislators began to push for more structure, particularly at the state level. The Summer 2026 AI Safety Index highlights the urgent need for enhanced safety frameworks in the face of these risks.
AI safety has largely relied on voluntary norms and fragmented guidance, now giving way to emerging state oversight
California and New York have already adopted laws that require frontier developers to document how they manage catastrophic risks and to publish safety and transparency reports for powerful systems. These state measures mark a first attempt to define obligations for developers whose systems could plausibly cause systemic harms, but they operate within individual jurisdictions and rely heavily on self-reporting and audits.
Anthropic has been an active participant in this shift, backing state bills that require external evaluations of safety protocols and more detailed disclosure of risk management practices. In its recent policy work, the company now argues that the next step should be federal rules with real enforcement powers, modeled loosely on how aviation regulators certify aircraft before they carry passengers.
Inside Anthropic’s Advanced AI Framework
Anthropic’s Advanced AI Framework is an attempt to spell out what binding frontier AI regulation could look like in practice, with concrete thresholds and enforceable duties rather than generalized principles. At its core is a simple idea: the most capable models must undergo mandatory independent safety tests, and regulators must have the authority to block or deter deployment when those tests uncover unacceptable risk. The framework reflects a wider recognition that government regulation is needed to manage escalating catastrophic AI risks as capabilities grow.
The framework restricts these obligations to what it calls covered developers, meaning organizations that build the most advanced models and clear both a compute threshold and a revenue or spending threshold. The technical trigger is training a model using more than ten to the power of twenty-five floating point operations, which corresponds to the high end of current frontier scale. On the financial side, a developer must either earn more than five hundred million dollars in AI-related revenue each year or spend more than one billion dollars annually on AI research and development.
In other words, the framework deliberately concentrates regulatory burdens on a small group of highly resourced frontier labs and large technology firms, while leaving smaller developers and lower compute open weight projects outside the mandatory regime. This reflects a strategic choice to target the systems most likely to create catastrophic or systemic risks, rather than imposing broad compliance costs across the entire ecosystem of everyday AI tools.
For each covered model, developers would be required to test capabilities against four enumerated risk categories that focus on catastrophic or systemic harms. Public reporting indicates those categories include biological threats, cybersecurity harms, loss of control over advanced systems and automated AI research and development that could accelerate technological progress in ways that outstrip traditional oversight. Developers would need to document how identified risks are mitigated and show their risk management strategies before deployment.
The framework goes beyond testing by requiring a robust security program for model weights and training infrastructure, recognizing that powerful models and their parameters are attractive targets for well-resourced attackers. It also calls for regular system cards, safety frameworks and risk reports that explain how catastrophic risks are assessed and mitigated, all of which must be submitted to at least one qualified independent evaluator. That evaluator is expected to publish technical findings that inform regulators, policymakers and downstream users, creating a public record of model capabilities and risk posture.
Enforcement is anchored in civil penalties calibrated to a developer’s global annual revenue, with escalating consequences for repeated violations or material misrepresentation. This is meant to give regulators genuine leverage to block or deter deployment of dangerous frontier models and to ensure that safety obligations are not merely aspirational.
Federal preemption and the state laboratory
Anthropic is not calling for federal rules in a vacuum. The company explicitly links its support for federal preemption of state AI laws to the passage of a stringent national statute focused on catastrophic risks. The argument is that Congress should override state experiments only if the federal framework provides clear authority and real tools to prevent deployment of dangerous frontier systems.
This position acknowledges the role states have played as early laboratories for AI oversight. California in particular has advanced bills that require mandatory safety testing and protective protocols for models that exceed defined cost or computational thresholds, reinforcing the idea that the riskiest systems deserve special scrutiny. Under these laws, developers of powerful models must publish safety frameworks, maintain security and shutdown mechanisms and produce audit reports that detail compliance and remaining gaps.
Anthropic’s framework aligns with this threshold-based approach but tries to standardize it at the federal level, so that frontier developers face a coherent set of expectations rather than a patchwork of state mandates. At the same time, the company appears wary of a federal vacuum in which preemption wipes away state protections without replacing them with something stronger, a scenario that would undermine current progress in places like California and New York.
Why open weight models are part of the story
A notable feature of Anthropic’s proposal is that it explicitly includes powerful open weight models within the scope of mandatory testing and security requirements. Open weight systems are those whose parameters can be widely copied, modified and integrated into new applications, which can dramatically expand both positive innovation and potential misuse once they are released.
By insisting that the most capable open weight models undergo independent safety evaluation and adhere to strict security practices, Anthropic is effectively arguing that openness does not exempt a developer from responsibility for catastrophic risks. The concern is that unrestricted access to highly capable models could enable actors with minimal resources to exploit dangerous capabilities, such as advanced biological design or highly automated cyber intrusion tools. The framework therefore treats powerful open weight releases as frontier deployments that must be justified through rigorous safety evidence and ongoing monitoring.
Implications for technology and business
For large AI labs and major technology platforms, this framework would mean that frontier model development and release become more like regulated infrastructure projects than typical software launches. Model training plans would need to account for the compute threshold, while business units would have to prepare for independent assessments, public system cards, detailed risk reporting and potential regulatory delays if safety tests surface serious concerns.
The financial penalties tied to global revenue create meaningful downside risk for cutting corners on safety, particularly for firms with multibillion-dollar AI businesses. Over time, this could push boards and investors to treat catastrophic risk as a core governance issue, rather than a specialized technical topic, which would influence hiring, internal audit and enterprise risk management practices.
At the same time, the narrow scope of the framework means that most startups and smaller companies deploying AI through application programming interfaces or lower compute models would not face direct regulatory obligations, at least initially. Many of these firms may still choose to adopt elements of frontier safety practices voluntarily, both to reassure customers and to prepare for the possibility that thresholds evolve as models become more capable.
For the broader economy and labor market, Anthropic connects its safety agenda to concerns about AI-driven job losses and structural unemployment. The company has argued that robust federal standards are part of a wider response that should also include modernizing unemployment infrastructure and social support systems, so that workers can adapt to rapid changes in demand for different skills. This framing treats safety and economic resilience as intertwined, rather than separate policy tracks.
Societal risks and governance challenges
From a societal perspective, the Advanced AI Framework reflects growing recognition that frontier models pose qualitatively different risks from conventional automation tools. Testing for biological misuse, cyber escalation and loss of control acknowledges that some capabilities could enable actors to cause widespread harm without traditional state capacity.
However, implementing this framework raises several hard questions. Measuring catastrophic risk in a model that has never been deployed at scale is inherently uncertain, and independent evaluators will need access, expertise and standardized methods to make credible judgments. Governments will have to decide what level of risk justifies blocking deployment, and how to remain transparent and accountable when those decisions affect national competitiveness in AI.
There is also the issue of global coordination. Frontier AI development is not confined to a single country, and models trained outside the United States could still be used by domestic actors. A unilateral US framework may encourage other jurisdictions to adopt similar rules, but it could also create tensions if different regulators reach different conclusions about what counts as acceptable risk.
Finally, there is a debate about the balance between caution and innovation. Mandatory testing and the threat of deployment blocks could slow down the release of some models, which some industry voices may see as a competitive disadvantage. On the other hand, clear safety rules and predictable governance can give companies and researchers confidence to invest in long-term projects that rely on frontier capabilities, knowing that the surrounding ecosystem is not purely reactive.
Takeaways and what to watch next
Anthropic’s push for mandatory independent safety tests on the most capable AI models is a significant step in the evolution of AI governance, moving from voluntary commitments to proposed federal rules with teeth. By combining a concrete compute threshold with financial criteria, the framework focuses attention on a small set of frontier developers whose systems are most likely to create catastrophic or systemic risks.
The inclusion of powerful open weight models, the emphasis on independent evaluation and public reporting and the linkage to economic resilience all point to a comprehensive view of what responsible frontier development should look like. At the same time, the debate over federal preemption, the role of state laws and the practical challenges of risk measurement and enforcement are far from resolved.
Over the next few years, the most important signals to watch will be whether Congress chooses to codify something like Anthropic’s Advanced AI Framework, how regulators interpret their authority to block or deter deployments and how other major labs respond in their own policies and releases. The choices made now will shape not only the safety profile of frontier AI, but also the pace and direction of innovation across the entire ecosystem.
Conclusion
Anthropic’s push to make safety testing mandatory for powerful open weight artificial intelligence models is a sign that frontier systems are starting to be treated less like software experiments and more like critical infrastructure. The debate matters right now because governments are tightening export controls, early safety laws are coming into force, and open models are advancing fast enough that a single release can reshape both innovation and national security policy in a matter of months.
How the debate over open weight models reached this point
For most of the past decade, the story of artificial intelligence has been a tug of war between open publication and closed development. Early breakthroughs in deep learning spread through open papers and public code, and that openness helped a relatively small research community grow into a global ecosystem.
As capabilities improved, models became more obviously dual use. The same system that writes helpful code can also generate malware, and the same model that assists with biology research can lower the barrier to designing dangerous pathogens. Governments started to pay attention to how model weights are distributed because once weights are widely available, control over the system effectively disappears.
In two important ways, policy has so far treated open and closed models very differently. First, recent United States export controls focus on very capable closed model weights that require a license to export when they are trained above a high compute threshold, while open weight models fall largely outside these controls even when they approach similar capability. Second, many early safety and transparency rules, such as the California frontier model law, are written for a handful of large developers and do not yet directly constrain most open releases from smaller labs and academic groups.
At the same time, policymakers have been cautious about moving too quickly against open models. A major report from the United States National Telecommunications and Information Administration in twenty twenty four explicitly concluded that evidence was not yet sufficient to justify broad restrictions on dual use foundation models with widely available weights. Instead, the report recommended that government agencies focus on monitoring risks, supporting external research, developing benchmarks and definitions, and encouraging or compelling audits and transparency before considering stronger interventions.
Within industry, coalitions such as the Frontier Model Forum have been building the technical foundations for serious safety evaluations. Their work on third party assessments and taxonomies for biological risk testing is meant to give both companies and regulators a menu of evaluation methods that go beyond basic benchmarks. Anthropic’s own responsible scaling policy has evolved in that context, using adversarial testing and red teaming to track whether models are approaching capability thresholds that could pose catastrophic risks.
That is the backdrop for Anthropic’s new push on mandatory testing for open weight systems. It is less an isolated proposal and more a bid to connect a number of parallel efforts into a coherent regime.
What Anthropic is actually asking for
In a detailed position statement on open weight models, Anthropic argues that all sufficiently capable models should undergo mandatory safety testing, regardless of whether their weights are open or closed. The company stresses that it is not calling for a blanket ban on open weight models and says explicitly that it does not support such a ban. Instead, the emphasis is on three levers.
The first lever is control of advanced chips, in particular keeping large quantities of the hardware needed to train frontier models out of the hands of actors that are likely to ignore safety norms. The second is enforcement against industrial scale distillation and copying, where an adversary uses another company’s closed model as a shortcut to train a high capability open model. The third lever, and the one reshaping the current debate, is a requirement that all sufficiently capable models undergo standardized safety testing by independent third parties before release.
Anthropic’s comment to the same United States agency that produced the open weight report makes that last point particularly clear. The company calls for a standardized safety testing regime for powerful general purpose models, open and proprietary, carried out by independent third party evaluators rather than left solely to internal company teams. The tests, in this view, become a prerequisite for deployment in the way that safety certifications are prerequisites for operating aircraft or nuclear plants.
The internal practices Anthropic highlights for its own systems illustrate what that regime might look like. The company’s responsible scaling policy sets out capability thresholds tied to catastrophic risks, and it requires comprehensive testing when models approach those thresholds. Its transparency materials describe adversarial evaluations for cyber security, biological misuse and other sensitive capabilities, along with systematic red teaming and bug bounty programs aimed at uncovering unexpected behaviors.
Taken together, the message is not that open models are uniquely dangerous, but that capability is what should trigger scrutiny. Once systems cross a certain capability line, Anthropic argues, they should face the same minimum safety bar whether they are closed, licensed to a few partners, or fully open weight.
How mandatory testing could work in practice
No government has yet written a full legal blueprint for the kind of regime Anthropic is urging, but existing rules and industry practices give some hints. The California frontier model law requires large developers to publish company wide safety and risk management plans, align those plans with national and international standards, include independent assessments, and release transparency reports that summarize catastrophic risk evaluations before major model releases. Although the law currently applies to only a small set of companies, it shows what a pre deployment documentation and testing obligation can look like in practice.
The Frontier Model Forum’s work on third party assessments and biological safety evaluations gives a technical starting point. These efforts describe how models can be evaluated for high risk biological assistance, what kinds of test environments and safeguards are needed, and how to combine benchmark style tests with more open ended probing for emergent capabilities. Anthropic’s own internal policies go a step further by tying evaluation cadence and depth to effective compute and other signals that a model may be approaching dangerous capability thresholds.
A mandatory regime along these lines would likely involve several elements. Capability thresholds could be defined using a mix of training compute, performance on sensitive evaluations, and qualitative judgments about autonomy and generality. Independent labs or accredited auditors would run agreed test suites for cyber, biological and other high risk domains under controlled conditions, report results to both the developer and regulators, and publish summaries that downstream users can consult. Governments might require enhanced post deployment monitoring and incident reporting for any model that passes through this process, echoing the way California obliges frontier developers to disclose serious safety incidents within strict time limits.
Importantly, Anthropic’s public statements emphasize exemptions for less capable models from startups and academia. The intent is to avoid choking off ordinary research and product development, while ensuring that models which could plausibly contribute to catastrophic harm face rigorous review. This choice mirrors the tiered approach in export controls, where only models above certain capability or compute thresholds are subject to licensing, although current rules mostly focus on closed frontier weights rather than powerful open ones.
What this means for open weight models
For the open model ecosystem, the core shift is conceptual. Today many rules implicitly treat open distribution of weights as something like a categorical dividing line. Open weight models often fall outside export controls, and policymakers remain wary of restricting them without strong evidence of concrete harm. Anthropic’s proposal instead says that open and closed models should be treated similarly once they reach a certain level of capability.
In practice, that would mean that the most powerful open weight releases could face longer lead times and higher compliance costs. A lab that wants to publish a high end language or multimodal model with widely available weights might first need to coordinate with independent evaluators, pass cyber and biological misuse tests, implement security controls around training data and internal tools, and demonstrate plans for monitoring downstream abuse. For small teams working on modest models, little would change. For the developers pushing the frontier of what open models can do, the bar would rise significantly.
The upside for open projects is that a credible testing regime could make it easier to defend openness in public debate. If a developer can point to independent evaluations and a shared regulatory framework, the conversation shifts from moral arguments about openness versus secrecy to concrete questions about whether a specific model meets agreed safety standards. That in turn could give governments more confidence to allow some open releases that might otherwise be blocked entirely.
There are also competitive implications. Anthropic’s position is carefully framed as neutral between open and closed models once capability is taken into account. Nevertheless, large companies with existing compliance teams, legal departments and security infrastructure will find it easier to absorb mandatory testing than leaner open collectives. Over time, mandatory evaluations could become part of the fixed cost of playing at the very high end of open development, nudging the ecosystem toward fewer but more thoroughly tested open weight frontier systems.
Risks, criticisms and unresolved questions
Any move toward mandatory testing for open weight models will face pushback, and some of that criticism is well grounded. One concern is regulatory capture. If only a handful of large firms have the resources to navigate complex testing obligations, rules risk entrenching incumbents and weakening the independent open source projects that have historically driven much of the field’s progress. California’s law already applies to a small cluster of big developers and imposes extensive documentation and assessment requirements, and critics worry that further layering of obligations could unintentionally harden this small club as the only viable frontier actors.
Another concern is definitional. It is difficult to specify exactly which models are sufficiently capable to trigger mandatory testing, especially when capability does not scale linearly with training compute or parameter counts. Anthropic’s own policies attempt to solve this by focusing on capability thresholds tied to concrete risks, but that approach requires continual updating as both models and evaluation methods evolve. There is a real possibility that some dangerous capabilities will emerge between evaluation rounds or in deployment contexts that tests did not anticipate.
There is also the geopolitical layer. Export controls today draw a bright line between closed frontier model weights, which are tightly controlled, and open weights, which are largely exempt even when they are powerful. Anthropic’s proposal seeks to close the gap where a model is so capable that releasing it in open form without testing could meaningfully increase global security risks. Yet if testing regimes are imposed asymmetrically, they could simply push the most risky open development to jurisdictions with weaker rules, undermining both safety and the competitiveness of firms that comply.
Finally, there are questions about who runs the tests and who sets the thresholds. Independent third party evaluation sounds straightforward, but in practice evaluators will need access to model weights, sensitive internal tools and prompts that reveal dangerous capabilities, all of which raise security and intellectual property worries. Public bodies can fill some of that role, but many governments lack the technical capacity to design and interpret cutting edge tests on their own.
These are not arguments against any safety regime at all, but they do illustrate why the details matter as much as the principle.
How businesses and society should read this moment
For technology companies, the signal is clear. Safety evaluations are moving from voluntary best practice to likely regulatory requirement for frontier scale models, and that shift will not stop at closed systems. Anthropic’s call sits alongside emerging documentation and transparency duties in California, growing expectations from national security agencies, and an expanding body of industry best practices around third party assessments. Firms that plan to train or deploy high capability models with open weights should start treating safety testing as a core part of the product lifecycle rather than a last minute checklist.
For policymakers, the proposal offers both a roadmap and a warning. It shows that at least some frontier labs are ready to accept binding obligations, including on their own future open releases, in exchange for a clearer and more predictable regulatory environment. At the same time, it highlights the need to design rules that preserve room for open scientific work, support smaller innovators, and recognize the global nature of model development. Aligning export controls, safety testing and transparency regimes across jurisdictions will be difficult, but without that alignment, the most risky actors will seek out the weakest link.
For the broader public, the debate is a reminder that the way frontier models are governed is still very much in flux. The decisions made over the next few years about testing, openness and accountability will shape not only how safe powerful models are, but also who gets to build them and on what terms.
The road ahead
Anthropic’s push for mandatory safety tests on powerful open weight models is an attempt to move the conversation beyond a simple choice between open and closed systems. It frames rigorous third party evaluations as the neutral infrastructure that allows societies to enjoy the benefits of advanced models without sleepwalking into catastrophic risks. Done well, such a regime could create a more level field between open and closed development, give policymakers better tools for risk management, and preserve much of the dynamism that open models have brought to the field.
Done poorly, it could harden the dominance of a few large firms, slow down beneficial open research, and give a false sense of security if tests fail to keep pace with rapidly evolving capabilities. The coming years will show whether governments and industry can translate the principle of mandatory testing into a practical system that is robust, fair and adaptable. That outcome will matter at least as much as any single model release in determining whether advanced artificial intelligence becomes a broadly trusted tool or a persistent source of systemic risk.









1 comment
Comments are closed.