JOURNAL / 2026.09.11
Anthropic's threat report shows how AI is lowering the cost of malicious operations
Anthropic's cyberattack, surveillance, weapons, and biology cases move the debate from hypothetical capability to observed use, but they remain a selected sample that is difficult to audit externally.
The most important finding in Anthropic's new threat report is not that a model invented an unknown form of harm. It is almost the opposite: human actors are inserting AI into operations that already existed—computer intrusion, surveillance, propaganda, fraud, weapons development, and dual-use science—to do more work with fewer specialists and in less time.
Detecting and countering misuse of AI: September 2026, published on September 10, brings together cases the company says it detected and disrupted between December 2025 and August 2026. It covers seven harm categories and actors ranging from individual criminals to groups allegedly linked to states. This is a material expansion from the November 2025 report, which focused on one highly automated cyberespionage campaign.
It is also worth establishing at the outset what this document is not. Anthropic acknowledges that it selected its "most notable and novel" cases, not a representative sample. It publishes no denominator from which to calculate what share of Claude use is malicious, nor the complete records that would let a third party reproduce its attributions. Apart from one illicit capability-extraction campaign, the cases did not use the newest Fable or Mythos models either. The relevant signal is not that only the frontier matters, but that already diffused capabilities are enough to change the cost of an operation.
The novelty is in the workflow
In the cyber cases, the routes in remained familiar: stolen credentials, exposed services, unpatched systems, and deception. What changed, according to Anthropic's records, was the amount of labor one operator could delegate. Agents wrote tools, explored systems, processed data, and handled several victims in parallel. In one campaign, the company says agents did nearly all the work; in others, a person approved every consequential step.
That difference prevents an easy conclusion. Autonomy and severity are not the same variable. Humans retained target selection, monetization, and review of results, and some of the most serious compromises occurred under close human direction. AI's contribution was to compress time, cost, and required expertise. A conventional attack that was previously uneconomic may become viable even if the model does not independently choose whom to harm.
The surveillance cases make the same economics visible outside cyberattacks. Anthropic describes public offices and contractors using Claude to classify posts, compile dossiers, and build systems for tracking dissidents or religious communities. In one case, a single consultant developed a platform for Malian authorities; after the account was blocked, the system was ultimately deployed locally with another model. Cutting off provider access stopped development inside Claude, but it did not withdraw the artifact or remove the demand.
The biology section calls for even greater care. The company presents five uses that could support dangerous research, ranging from experimental planning to proposal writing and computational redesign of dual-use molecules. It does not claim that Claude enabled anyone to build a biological weapon or caused a measured capability uplift. In fact, it concludes that the cases do not establish an imminent biological threat. They do show two complementary limits: filters blocked or degraded several recognizable requests, while scientifically legitimate and potentially harmful work could share methods, language, and intermediate goals.
Searching for a forbidden word is not enough in that setting. Anthropic proposes combining filters with trusted-access programs, institutional signals, identity verification, and data retention. That is an understandable response to the problem it observes, but not a neutral solution: it moves decisions about who may conduct research, what activity is suspicious, and how much must be retained for monitoring toward a private company.
The platform sees earlier; it does not see everything
A model provider occupies a singular observation point. It may detect a campaign while it is being written or programmed, before its results appear on a social network, a compromised system, or in a laboratory. It can connect accounts, block them, and turn new patterns into controls. Anthropic says it shared indicators with authorities, victims, and other platforms where appropriate.
But its visibility drops precisely where much of the harm begins. For influence operations, the report itself explains that it needs open-source research and data from other platforms to learn whether content reached real people. In surveillance and malicious software, a user can migrate to another service or locally run what has already been built. And for the public reader, the most sensitive evidence remains mediated by the provider's selection, attribution, and redactions. Forbidden Stories' investigation of S2T corroborates that a surveillance product under that name was marketed to armed forces as early as 2023; it does not independently verify the later role of Claude that Anthropic attributes to its records.
Anthropic's new evaluations add a second kind of evidence. Models are improving at locating photographs or text and at writing controllers for simulated drones. These tests are useful for observing a trajectory, not for directly measuring real-world harm: the human geolocation comparison comes from a different task with different imagery and conditions, and a simulator does not reproduce all the engineering or friction of a physical weapon. The evidence is strongest where the two lines meet: evaluations show tasks a model can make cheaper, while records show actors trying to incorporate them into real operations.
My reading is that this report shifts the center of gravity in the malicious-use debate. It is no longer enough to ask whether a model would answer a dangerous request in a test. We have to examine the entire system: accounts and resellers, connected tools, credentials, parallel execution, human decisions, off-platform deployment, and the response of whoever can observe each segment.
That requires more than classifiers. Keys that turn a legitimate account into someone else's infrastructure must be protected; useful indicators should be shared without publishing attack recipes; migration to other models should be measured; and provider attributions should face independent review using protected data. Denominators and selection criteria should also be published when safe: a collection of extreme cases warns about possibilities, but it cannot estimate prevalence, trend, or the overall effectiveness of blocking.
The difficult tension remains. Greater observability may improve detection while expanding private surveillance over legitimate conversations. Restricting closed models can raise the cost of a campaign without preventing it from continuing with local weights. As researcher John Thickstun told the Associated Press, leaving these decisions solely to companies amounts to delegating society-scale value judgments without sufficient democratic deliberation.
The report provides serious evidence that AI is already reducing the operational cost of harm; it does not demonstrate that every threat originated with AI or that models act without human direction. Its practical consequence is more sober and perhaps more important: defense must adapt to adversaries able to multiply familiar work, and society needs an auditable way to govern the extraordinary visibility accumulating with providers.
Sources
- Anthropic, Detecting and countering misuse of AI: September 2026, September 10, 2026.
- Anthropic Frontier Red Team, Measuring tactical intelligence targeting and conventional weapons capabilities of AI models, September 10, 2026.
- Anthropic, Disrupting the first reported AI-orchestrated cyber espionage campaign, November 13, 2025.
- Forbidden Stories, When your “friends” spy on you: The firm pitching Orwellian social media surveillance to militaries, February 20, 2023.
- Associated Press, Anthropic says it blocked misuse of its AI that could have supported biological weapons, September 10, 2026.