JOURNAL / 2026.08.06
US completes a nonpublic framework reviewing closed—but not open-weight—AI models
The voluntary mechanism opens up to 30 days of pre-release testing for certain closed models, but its rules are not public and open weights sit outside that advance review.
The United States now has a process through which the government can examine some AI models before they reach third parties. What the public does not have is the process itself to read.
On August 3, the White House confirmed that it had completed the voluntary framework ordered two months earlier by its deadline. The June 2 executive order is public: it directs officials to create classified tests of advanced cyber capability, use them to identify “covered frontier models,” and let developers give the government access for up to 30 days before providing those models to other trusted partners. It also makes clear that the mechanism does not create a mandatory license, preclearance, or permit to release a model.
The document that turns that outline into an operating process has not been published. According to the confirmation reported by Axios, the administration considers the framework complete, is discussing next steps with industry, and does not plan to release it. The capability threshold was already required to be classified under the order; the rest of the framework did not carry that explicit designation.
A meeting with companies on August 4 revealed a second important distinction. Sources briefed on the contents said a covered model must be closed, state of the art, and pose national-security risks. Open-weight models are excluded, while the concrete definitions of “state of the art” and “national-security risk” have not been made public. These are consistent reported details about a withheld text, not a rule the public can inspect directly.
That asymmetry deserves more attention than the word “voluntary.” The United States has created a pre-release review for the kind of model whose provider retains access control, while this mechanism offers no review before publishing files that anyone will be able to download, modify, and redeploy.
Two different times to evaluate
Excluding open weights from this framework does not mean the US government has stopped testing them. NIST’s Center for AI Standards and Innovation (CAISI) lists unclassified evaluation of demonstrable cybersecurity, biosecurity, and chemical-weapons risks among its functions. This year it has published assessments of DeepSeek V4, GLM-5.2, and Kimi K3.
The GLM-5.2 assessment shows the temporal difference. Z.ai released the model with open weights on June 16; CAISI completed its evaluation on July 8 and published the findings nine days later. The center could measure general capability, cyber capability, and safeguards without a confidential window granted by the developer. It also warned that, however robust refusals may look in the tested version, a person hosting the weights can circumvent them.
That later access has a scientific advantage: an external evaluator can retain the artifacts, change the configuration, remove serving layers, and repeat tests. With a closed model, the evaluator depends on the interface and conditions the provider maintains. But downloading does not rewind the clock. If a model crosses a dangerous threshold, the first public measurement arrives after the files may have been copied. A later assessment can inform defenses, procurement, or use restrictions; it cannot make the initial release reversible.
Pre-release review of a closed model has the opposite property. It allows testing before distribution, while capabilities or safeguards can still be modified, and supports agreement over who receives early access. In exchange, the evidence stays inside a confidential relationship between the developer and the state. The provider may participate or decline, the executive order does not require publication of the results, and the public still does not know what finding would delay a release, narrow access, or merely produce a recommendation.
These are not two versions of the same test. They are different instruments with different possibilities and failure modes:
- a closed model can be examined early and corrected from a single control point, but the evaluator receives the system the provider offers and the conclusions may never leave the room;
- open weights can be studied more independently and by many parties, but a test of the original version does not certify its derivatives, and any warning may arrive after irreversible distribution.
The US decision does not resolve that tension. It chooses advance cooperation for one branch and later observation for the other.
A secret benchmark does not require an illegible policy
There is a legitimate reason not to publish the exact tasks in an evaluation. If a lab can train against them, they stop measuring general capability and start measuring familiarity with the exam. CAISI already combines public sets with held-out tests and explains at least part of its harness, budget, aggregation, and limitations. Its DeepSeek V4 evaluation, for example, identifies the domains measured, how the model was served, and where a comparison was not equivalent.
Protecting test items does not require hiding all the governance. Officials could publish the institutional definition of a covered model; who decides and what review is available; which version of a system is tested; how tools, safeguards, and budgets are recorded; which classes of result trigger action; how a disagreement is documented; and what aggregate summary will appear after release. None of that teaches a model to memorize a cyber test.
Without that public layer, there are practical consequences. A company building on an API cannot interpret the provider’s participation as a safety seal because it does not know the result—or even whether the tested version matches the one it uses. A researcher cannot tell whether two labs received the same protocol. A smaller developer cannot predict when it will be considered frontier except through private government discussions. International partners see the mechanism’s boundaries only through leaks and official remarks.
My reading is that the framework can produce useful security work and still be weak public policy. A classified evaluation may discover capabilities that should not be taught; an opaque process makes it impossible to verify whether that discovery changed a decision. Test confidentiality and procedural accountability are not opposites.
The most revealing fact in the next model launches will not be a company saying it “worked with the government.” Watch whether it identifies the evaluated version, explains what changed after review, publishes an adequate technical card, and whether the government provides comparable aggregate findings. For open weights, the additional measure is how long independent evaluation takes after release and whether it covers both the base capability and the real systems that turn it into actions.
The new framework changes who can look at a closed model before launch. It still does not let us know what they saw, what they decided, or why. Until that minimum traceability exists, this is not a public approval. It is a confidential cooperation channel whose effectiveness remains a claim to be tested.
Sources
- White House, Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security, June 2, 2026.
- Axios, White House finalizes AI framework behind closed doors, August 3, 2026.
- Axios, Inside Trump's AI framework, August 4, 2026.
- NIST, Center for AI Standards and Innovation (CAISI), accessed August 6, 2026.
- NIST, GLM-5.2 assessment, July 17, 2026, and DeepSeek V4 Pro evaluation, May 1, 2026.