TL;DR
The US government has mandated a classified benchmarking process for advanced AI models, due by August 1, 2026, involving NSA, Treasury, and other agencies. Participation is voluntary but may influence federal procurement and industry standards.
On June 2, the Biden administration announced that by August 1, 2026, the NSA, Treasury, and other agencies will establish a classified benchmarking process to evaluate the cyber capabilities of advanced AI models. This process will define thresholds for what constitutes a covered frontier model and will influence federal procurement and security policies, marking a significant shift in US AI oversight.
The executive order, signed by President Trump on June 2, mandates the creation of a classified cyber-capability benchmark for AI models, with the NSA Director making the final designation. Alongside this, a voluntary framework will allow developers to share models with the government for up to 30 days before public release, enabling pre-deployment assessments. This framework aims to foster collaboration while maintaining confidentiality of sensitive evaluation criteria.
Additionally, the order establishes an AI cybersecurity clearinghouse under Treasury to share vulnerability intelligence between industry and critical infrastructure operators, and allocates funding for AI vulnerability detection tools and federal cyber talent recruitment. The process is voluntary but may carry significant industry implications, as participating vendors could gain preferred status in federal procurement, effectively making opt-in a de facto requirement.
Implications of Classified AI Benchmarks on Industry and Security
This development signals a major shift in US AI governance, moving from a hands-off approach to active oversight involving classified assessments. The establishment of a classified benchmark introduces a new layer of security evaluation but raises concerns about transparency, potential bias, and the difficulty of independent verification. Industry stakeholders face strategic decisions about participation, which could influence market access and government contracts. For national security, the process aims to better identify and mitigate AI vulnerabilities, but the secrecy surrounding benchmarks may complicate research and international cooperation.

The Developer's Playbook for Large Language Model Security: Building Secure AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Evolution of US AI Security Policies
The order reflects a response to growing concerns about AI capabilities and vulnerabilities, especially in cybersecurity. Previously, the US government relied on voluntary cooperation and public standards, such as the EU AI Act’s transparency requirements. An earlier version of this executive order was reportedly withdrawn over fears it might hinder US competitiveness. The current framework emphasizes classified evaluations, aligning AI security with traditional defense assessment practices, and marks a notable shift from prior policy that largely avoided direct oversight.
“The classified benchmark will serve as a key tool for assessing AI cyber capabilities, with the NSA making final designations based on sensitive criteria.”
— Official familiar with the order
AI cybersecurity monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Uncertainties Surrounding Classification and Industry Impact
It remains unclear how the classified benchmarks will be formulated, whether they will be challenged or reviewed publicly, and how precisely participation will influence industry access to federal markets. The scope of government evaluation—whether it includes proprietary data or IP protections—is also still under discussion. Additionally, the long-term effects of a classified system versus transparent standards are yet to be determined, especially regarding international cooperation and research transparency.

TANGEM Crypto Wallet Pack of 2 – Trusted Cold Storage Hardware Wallet
- Proven Security: Military-grade EAL6+ security with no hacks
- Easy Blockchain Access: Manage 90 blockchains with one tap
- Wide Cryptocurrency Support: Access 14,100+ coins, tokens, NFTs, DeFi
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Industry and Government Stakeholders
Industry players must decide whether to participate in the voluntary pre-release framework before August 1, 2026. Developers and vendors will likely weigh the benefits of trusted partner status against the risks of sharing sensitive models and data. Meanwhile, Congress and oversight bodies may debate whether to move from voluntary to mandatory testing requirements, potentially transforming the framework into a more stringent approval process. The NSA and Treasury are expected to finalize the benchmark criteria and operational procedures by the August deadline, with ongoing discussions about transparency and international implications.

AI Engineering: Building Applications with Foundation Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the purpose of the classified benchmark for AI models?
The benchmark aims to evaluate the cyber capabilities of advanced AI models to inform security and procurement decisions, with thresholds defining when a model is considered a covered frontier model.
Will participation in the pre-release review be mandatory?
No, participation is currently voluntary, but industry experts suggest it could become a de facto requirement for federal contracts, depending on how the framework develops.
How does this order compare to European AI regulations?
Unlike the EU AI Act, which sets public, contestable thresholds based on compute power, the US order establishes a classified, non-public benchmark, raising concerns about transparency and independent verification.
What are the potential risks of classified benchmarks?
Classified benchmarks could drift from original intent, encode vendor-favorable assumptions, or be wrong, with no external review or falsification possible, potentially impacting industry trust and research transparency.
What happens if a developer refuses to participate?
Refusing to participate may limit access to trusted partner status and federal contracts, potentially affecting market opportunities and industry standing, but the order does not mandate participation.
Source: ThorstenMeyerAI.com