Anthropic says its Claude models can orchestrate much of an early drug-discovery protein-design workflow, according to The Decoder’s report on two Anthropic experiments. The company tested Claude on de novo minibinder design: creating small proteins intended to attach tightly to target proteins and potentially block or alter their function. The results are notable, but still early. The Decoder reports that Anthropic says the work beat typical industry hit rates, while also noting that an independent review of the results is still pending. That matters because the central claim is not just that a model suggested designs, but that language-model agents managed a full stack of specialized tools and produced candidates that were later tested in the lab. In the minibinder experiment, Anthropic used Mythos Preview and Opus 4.8 against 16 target proteins, according to The Decoder. Fifteen targets produced usable measurements, and Claude-generated designs succeeded on 14 of those 15. Across 1,320 lab-tested designs, 354 bound to their targets, giving a 26.8% hit rate. The Decoder also reports that when looking only at designs Claude ranked first, 49% bound. The strongest reported number came with an important caveat. In a multi-target mode, where targets were handled together within 48 hours, Mythos Preview reached a 26.7% hit rate and Opus 4.8 reached 22.6%, The Decoder reports. When Mythos Preview worked on each target individually, the hit rate rose to 35.1%, but with 2.8 times the compute budget per target. Anthropic’s authors acknowledge, according to the report, that focus and budget cannot be separated in that comparison. Anthropic’s claimed benchmark is the current typical range of 10% to 15%, which The Decoder says Anthropic drew from publicly documented campaigns in the proteinbase.com database. That makes the headline result potentially important if reproduced, but the comparison should be read as Anthropic’s framing rather than an independently validated standard within this cluster. The technical mechanism is also important: The Decoder reports that Anthropic did not build a new protein model. Instead, Claude installed and ran open-source specialty software already used in the field. Backbone generation came from tools including PXDesign, RFdiffusion3, Genie 3, FreeBindCraft, BoltzGen, RFdiffusion, and Proteina-Complexa, among others. The amino acid sequences were mostly computed with SolubleMPNN, a variant of ProteinMPNN. For filtering and ranking, Claude used ESMFold2, ESMFold2-Fast, and Protenix v2, according to The Decoder. These programs predict binder-target folding and provide confidence scores for whether binding is likely to work. The report says AlphaFold-3 weights, Rosetta, and ESM3 were excluded for licensing reasons. The agent setup depended on a protocol prompt of about 16,000 words. The Decoder reports that about one third of the prompt covered scientific guidance and a reading list, while the rest addressed scheduling, delegation to sub-agents, verification, and budget discipline. The prompt did not specify the epitope — the target surface region to attack — for any target. Humans still bounded and interpreted the process. According to The Decoder, people chose the targets, wrote the prompt, ordered synthesis, and interpreted the measurement data. The compute ran through Modal, with budgets of $50,000 per multi-target campaign and $10,000 per single target. In between, Anthropic’s report says only short, non-technical instructions were needed to resume after infrastructure outages. Who benefits: Labs already using open-source protein-design tools could benefit if agent orchestration makes those workflows easier to run and repeat. Cloud execution platforms such as Modal are also part of the reported workflow, though the cluster does not establish any commercial impact. Who's exposed: Teams that rely on manual orchestration of specialized protein-design software may face pressure to evaluate agent-based workflows if the results are independently reproduced. The evidence is not yet strong enough to claim a broad change in drug-discovery practice.