JOURNAL / 2026.09.13
Anthropic promises permanent outside evaluators and calls for pacing the AI frontier
Anthropic will give independent reviewers employee-like access and OpenAI promises to follow; it is a concrete oversight change, but no verifiable mechanism yet exists to measure or enforce the pace.
The most concrete part of Dario Amodei's call to slow AI is not the word “slow.” It is a door Anthropic promises to open: a permanent external team will be able to enter the company with access resembling that of its internal evaluators, verify safety practices, investigate incidents, and publish conclusions without Anthropic's editorial control.
The CEO made the commitment in a personal essay on September 12. He argues that capabilities are advancing faster than safety can absorb, driven in part by using AI to build the next generation of AI, and proposes “pacing the frontier” at three levels: embedded evaluators in each lab, coordination among democratic companies and governments, and global agreements that include China. Sam Altman responded publicly that OpenAI will also make the first commitment and shares the need to pace progress.
It is important to separate immediately what happened from what is desired. Anthropic has promised a new form of access, but it still has to invite the team and agree on a contract. OpenAI has announced that it will match the measure without yet publishing its design. There is no industry agreement to slow training, no shared threshold that compels one, and no international treaty. Today's verifiable change is narrower and perhaps more durable: two rival labs have publicly accepted that a one-off evaluation before release is no longer sufficient to oversee systems that also work during training and inside the company itself.
An audit that can inspect the process, not just the finished model
External evaluators already receive early model access. METR, the organization Amodei gives as an example, has worked with Anthropic, OpenAI, Google DeepMind, Meta, and Amazon. But those exercises are generally bounded in time and scope. In its May frontier risk report, METR was able to test internal models, question companies, and, at Anthropic, attack its monitoring system. Even so, each participant could silently leave before publication and control which nonpublic information appeared; the report acknowledged that its visibility into several companies' practices was limited.
The new promise tries to change that relationship. Amodei describes desks, badges, and company laptops; permissions and tools comparable to those of internal risk teams; direct conversations with employees; and a contractual right to publish findings about risks, incidents, practices, and denied access. Anthropic would retain narrow redactions for security, legal privilege, commercial sensitivity, or third-party confidentiality, but the reviewer could state that a redaction affected its conclusions.
That matters because it expands both the duration and the unit of evaluation. A pre-release test asks what a prepared version of a model can do under a particular harness. A continuing presence can also ask which model the lab actually uses internally, which permissions it receives, whether monitors cover every route, when a run is paused, how broken environments are filtered, and whether a fix remains active months later.
The evaluation incidents analyzed in this journal show why that continuity is material. The failures did not live only in a dangerous model response: they appeared in Internet egress, public packages, credentials, filters, and shared services. More recently, the GPT‑6 Astra release combined greater cyber capability with controls that depend on the interface and the observed trajectory. Reviewing only the checkpoint would miss much of the system that determines risk.
Access does not by itself mean independence
The proposal improves an obvious weakness in voluntary oversight: until now, the lab largely chose which model, information, and moment the reviewer saw. But “employee-like” does not yet define an audit. The participating organizations and their number, term, funding, appointment and dismissal process, reporting frequency, legal protection to investigate, and procedure for resolving a redaction dispute are all missing.
Authority is also missing. The evaluator will be able to observe and publish; the essay does not give it power to stop a training run or deployment. It can make visible that Anthropic is breaking a promise, but the consequence depends on management, customers, public opinion, or a regulator. That is an essential distinction between an independent second opinion and a public supervisor with a mandate defined in law.
Independence does not mean having no relationship with the company, either. It means making that relationship legible and preventing it from buying the outcome. METR says it does not accept payment for its evaluations, although companies provide access and tokens, and its report disclosed personal relationships and the incentive to preserve cordial collaboration. That transparency is useful. A permanent regime will also need stable funding outside the audited lab, multiple teams able to check one another, conflict-of-interest rules, and enough security that auditor access does not become another path to sensitive models, data, or vulnerabilities.
Anthropic already requires external reviews of some risk reports under its Responsible Scaling Policy. The proposed difference is moving from reviewing a document—even an unredacted one—to inspecting the processes that produce it. The jump will only be real if the final contract preserves discretion to pursue unexpected signals and if reports explain not only what reviewers were allowed to see, but also what remained outside their reach.
The verb “pace” still has no counter
The other two levels of the plan are far more ambitious and much less specified. Amodei proposes that labs in democracies agree on capability-linked safety requirements: if a model can defeat common isolation methods, for example, it would have to demonstrate alignment properties and sufficient controls before advancing. He also considers limits on compute, training runs, or internal use of AI to improve AI. Farther ahead, he imagines measures ranging from international bans on biological uses to a speed limit on recursive improvement.
The problem is not just securing cooperation; it is defining the quantity to be limited. Is it the compute in one run, observable capability, experiment volume, automated work, or progress across the whole lab? The research telemetry published by OpenAI has already shown that restricting one model class can move accelerators and people toward others. A verifiable brake needs to count substitution, post-training, agent fleets, and software improvements, not just a large named run.
Amodei's urgency should not be confused with timeline evidence, either. His concern that in six to twelve months a swarm could maintain a botnet capable of taking over “the entire internet” is an extrapolation, not the result of a published evaluation. Real incidents and partial R&D acceleration justify greater attention; they do not establish that capability or date precisely. OpenAI said this week that fully autonomous recursive self-improvement is not happening today, while acknowledging growing automation of parts of research and calling for rules to decide when development should slow or stop.
My reading is that the access commitment deserves more attention than the dramatic forecast. One lab has accepted that a third party should continuously observe how it turns policy into operations, and a leading competitor says it will do the same. If executed well, this could produce something missing from nearly every recent debate: contemporaneous evidence not selected solely by the company whose conduct is being assessed.
But the embedded evaluator is not yet the brake. It is the instrument that might tell us whether anyone is applying one. Moving the plan from statement to governance requires three separate things: measures that track real progress across the system, public criteria that trigger reduced activity, and an authority able to require it. External access strengthens the first; it may illuminate the second; it does not create the third.
The next test will be very prosaic. Anthropic and OpenAI will have to name the evaluators, publish the scope of their agreements, let them describe access limits, and accept adverse reports without turning every exception into commercial secrecy. If that happens, September 12 will have opened a new institution inside the frontier. If it does not, “pacing” will remain an intention that cannot be audited precisely when it matters most to know whether it is being honored.
Sources
- Dario Amodei, We Must Pace the Frontier, September 12, 2026.
- METR, Frontier Risk Report (February to March 2026), May 19, 2026.
- Anthropic, Responsible Scaling Policy, version 3.4 and updates accessed September 13, 2026.
- OpenAI, The AI policy window is open. We need to act, September 9, 2026.
- Axios, Anthropic, OpenAI CEOs call for slowdown in AI development, September 12, 2026; and Associated Press, coverage of the call and OpenAI's response, September 12, 2026.