JOURNAL / 2026.08.21

Claude designs 354 protein binders validated in the lab

Claude directed design campaigns against 16 targets and two labs measured every candidate; the result demonstrates molecular binding, not function or a finished drug.

A language model is no longer merely proposing a protein on a screen. In an experiment published by Anthropic, Claude directed design campaigns against 16 targets, delivered 30 candidates per target, and let physical synthesis and measurement decide the outcome. Of the 1,320 designs with interpretable data, 354 bound their targets. At least one succeeded for 14 of the 15 targets that could be evaluated.

This is unusually concrete evidence for an “AI for science” claim. It is also easy to describe badly. Claude did not generate the proteins from its own parameters, invent the protocol by itself, or produce 354 medicines. The contribution lies elsewhere: a general agent coordinated specialist tools, sustained decisions for one or two days, and delivered a complete portfolio to a lab without a person selecting favorable candidates after seeing the outcome.

The system was larger than Claude

Researchers selected the targets and wrote a protocol of roughly 30,000 tokens. They refined it during test campaigns and froze it before the reported experiments. They also supplied literature, internet access, connectors, GPUs, and budgets of up to $50,000 for a 48-hour multi-target campaign or $10,000 for a 24-hour campaign dedicated to one target.

Within that framework, Claude Opus 4.8 and an unreleased model called Mythos Preview researched each target protein, chose the surface to bind, installed and operated open-source structure-design, sequence-design, and co-folding models, iterated through filters and optimization, and ranked the final 30 candidates. Human operators approved access requests, resolved infrastructure failures, and placed the synthesis orders, but did not install the tools, choose binding surfaces, or alter the final rankings.

That is the practical change. Specialist models that propose and score structures already existed. The new work tests whether a general model can turn them into a long, coherent campaign without a specialist driving every step. It does not remove expertise: some of that knowledge is compressed into the protocol, corpus, and tools. It does move the bottleneck from knowing how to operate the entire computational stack toward specifying the goal, reviewing the process, and paying for physical verification.

From a human protocol to physical measurement: the Claude-directed protein-design system and the boundary of what it demonstrated

Physical measurement changes the quality of the claim

Anthropic did not test only a promising selection. Two contract research organizations, Adaptyv Bio and Twist Bioscience, received every delivered design. They produced and measured them separately using different surface-plasmon-resonance formats. Neither saw the other’s data or knew which model, campaign, or ranking position had produced a sequence.

One target aggregated and produced nonspecific measurements at both labs; its 120 candidates were excluded, leaving 1,320. To integrate the two readouts, the authors applied a fixed rule and released the raw signals, fits, and labels. The count of 354 denotes candidates with a concentration-dependent binding signal at one or both labs under that rule. It is not a simulated score. At the difficult end, none of the 90 designs against maltose-binding protein was confirmed; that failure is also in the public dataset.

That completeness matters more than the comparative headline. Anthropic reports a 26.8% overall hit rate and contrasts it with previous campaigns, but the technical report acknowledges that no human team worked with the same targets, tools, and budget. Results from four of the six competitions used for comparison were already available during design. The dedicated mode also received 2.8 times more compute per target than the multi-target mode, so the experiment cannot separate the effect of focus from the effect of budget. Hit rates are not fully comparable across publications either: Adaptyv explains that targets, assays, and even the definition of a “binder” vary.

The robust claim, then, is not that Claude defeated the best human designers. It is that Claude autonomously executed, at unusual scale, a complete process whose output underwent blind and exhaustive physical testing. The released dataset includes the protocols, per-design provenance, sequences, structure predictions, and laboratory measurements under open licenses. It makes the count auditable and the filters improvable, although it cannot exactly reproduce the agent because Mythos Preview and the Claude Science environment are not generally available.

Binding is not function

The most important boundary comes after the experiment. None of these proteins was tested for biological activity, and none had its structure solved experimentally. For five oligomeric targets, multivalency can strengthen the apparent affinity. Several sequences also derive from the same backbone and are not fully independent attempts. The authors themselves reduce the hit rate from 26.8% to 24.7% when retaining only the top-ranked sequence from each backbone.

A binding protein is a starting point, not a treatment. Researchers would still need to confirm where and how it binds, show that it causes the intended biological effect, study selectivity, stability, immunogenicity, distribution, and toxicity, and then pass through preclinical development and clinical trials. The result may compress an early stage that consumes weeks of computational work. It does not justify compressing the rest of the chain into the word “drug.”

Nor is this a capability ready for any developer today. Anthropic still keeps professional molecular-design requests outside general access to Fable 5, routing them to a less capable model while it develops trusted-access paths, according to its biology safeguards policy. Even with the agent, specialist compute, synthesis, and assays remain necessary. Those steps impose cost, but they are also checkpoints where goals, orders, and results can be reviewed.

My reading is that the most important part is not the 354 successes in isolation, but the link between digital autonomy and material verification. Many agents appear convincing because the environment that generates their work also grades it. Here, the molecules had to exist and produce a physical signal, and every attempt reached the scoreboard. That is a standard AI-enabled science should imitate.

The same experimental design shows where caution should concentrate. If general models learn to operate existing scientific tools, evaluating only what they can answer in text will no longer describe their real capability. Systems will need to be assessed with their tools, budgets, and access; human control over physical execution must remain meaningful; and failures need to be published alongside successes. This work does not demonstrate an autonomous drug factory. It demonstrates something earlier and important enough: a specialized part of molecular design can already become delegable, measurable, and much faster work.

Sources

← Back to journal