In two experiments, Anthropic put its Claude models to work on early-stage drug discovery. The company says its protein design results beat the usual industry hit rates. An independent review of the results is still pending.
Anthropic has published two experiments in which its Claude models took on tasks from the early phase of drug development. The first focused on designing so-called minibinders, small proteins that lock tightly onto a target protein and block or change its function. This principle underlies many drugs.
Designing such binders from scratch on a computer, rather than searching for them in nature, is called de novo design. According to the technical report, it still takes a series of expert decisions plus days of orchestrating specialized software and compute.
The models Mythos Preview and Opus 4.8 designed binders against 16 target proteins, 15 of which produced usable measurements. Claude succeeded on 14 of those 15 targets. Of 1,320 designs tested in the lab, 354 actually bound to their target, a hit rate of 26.8 percent. Looking only at the designs Claude ranked first on its own list, 49 percent bound.
In multi-target mode, where all targets were handled at once within 48 hours, the models reached 26.7 percent (Mythos Preview) and 22.6 percent (Opus 4.8). When Mythos Preview worked each target on its own, the rate rose to 35.1 percent, but with 2.8 times the compute budget per target. Focus and budget can't be separated, as the authors acknowledge.
For comparison, Anthropic cites today's typical range of 10 to 15 percent, drawn from publicly documented campaigns in the proteinbase.com database.
A language model runs a dozen specialized tools
Anthropic didn't build its own protein model. Claude installed and ran only open-source specialty software that the field already uses. The spatial scaffolds of the proteins, known in the jargon as backbones, came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), FreeBindCraft (135), BoltzGen (134), RFdiffusion (118), and Proteina-Complexa (100), among others. The amino acid chain that eventually forms this scaffold was mostly computed by SolubleMPNN, a variant of ProteinMPNN.
To filter and rank the candidates, Claude used a mix of ESMFold2, ESMFold2-Fast, and Protenix v2. These programs predict how binder and target fold together. They also give a confidence score for whether the binding should work. AlphaFold-3 weights, Rosetta, and ESM3 were excluded for licensing reasons.
The setup included a protocol prompt of about 16,000 words that ran as the system prompt in every agent. Only about a third of it is scientific guidance plus a reading list. The other two thirds cover scheduling, delegation to sub-agents, verification, and budget discipline. The prompt didn't specify which spot on the protein surface to attack, the so-called epitope, for any target.
The compute budget was $50,000 per multi-target campaign and $10,000 per single target, run through the cloud provider Modal. Humans picked the targets, wrote the prompt, ordered the synthesis, and interpreted the measurement data. In between, the report says, only short, non-technical instructions were needed to resume after infrastructure outages.
A contest winner beaten on the same assay plate
Validation was handled by the paid contract labs Adaptyv Bio and Twist Bioscience. They produced each design biologically, unchanged, and measured independently whether it binds to its target and how tightly. Binding strength is given as a KD value, measured in nanomolar (nM). The smaller the number, the tighter the binding. Small values matter for a drug because it then works at a lower dose.
Results varied wildly by target. For the immune receptor TREM2, 72 of 90 designs bound. For the growth factor VEGF-A, 54 of 90. The most telling case is RBX1, part of an enzyme complex that controls the targeted breakdown of other proteins in the cell. In an open design contest, only 9 of 245 newly designed candidates bound there. With Claude, it was 28 of 90. Anthropic had the contest's winning design rebuilt and checked on the same assay plate. It bound at 45 nM, while Claude's best design bound at 3.9 nM, roughly ten times tighter.
TNFα is considered especially hard. It's an immune system messenger that triggers inflammation and is the target of five approved biotech drugs, including Humira. Several earlier design approaches reported zero hits there. Claude produced 12 binders from 150 designs, all from Opus 4.8, none from Mythos Preview. But those twelve trace back to just four different scaffolds, so they're less independent of each other than the number suggests.
Binding to animal versions of the target proteins was only a secondary goal in the prompt, yet 130 of 233 tested binders also bound the mouse counterpart of their target. That's practically relevant because drugs are tested in animals before humans.
Two targets showed clear limits. BBF-14 is a computer-invented, barrel-shaped protein that doesn't occur in nature and therefore has no evolutionary history a design method could draw on. The bacterial maltose-binding protein MBP, in turn, has a smooth, water-loving surface that gives a binder little to grab onto. Against BBF-14, only three weakly binding designs worked, and against MBP, none of 90. Anthropic says the confidence scores from the folding prediction warned of neither failure. The designs against MBP and BBF-14 got almost the same scores as those against successful targets.
Chemistry analysis in under 25 minutes
In the second experiment, Opus 5 interpreted raw data from two standard measurements from a contract lab. Nuclear magnetic resonance spectroscopy (NMR) shows chemists whether they actually made the substance they meant to make. Liquid chromatography with mass spectrometry (LC-MS) reveals how pure the sample is. Both instruments output files in proprietary formats that are normally analyzed by hand in the manufacturer's software. From these raw files and prompts of one to three sentences, Claude delivered results in 23 and 19 minutes.
For the LC-MS file, the model found no suitable reader program and decoded the format itself. As a cross-check, it exactly reproduced the summary values the instrument had stored for all 2,664 measurement points. In the NMR analysis, Claude proposed the same follow-up experiment the lab had independently run three days after the first measurement, then corrected its own mistake. It first reported four missing signals, but the internal check found two. The much-cited purity values of 96.4 versus 96.33 percent rest on different baselines, according to the report. Calculated by the lab's method, Claude would land at 98.8 percent.
What's actually new here
Predicting protein structures and generating new proteins has been established since AlphaFold, RFdiffusion, and BindCraft. The actual design work here was done by these tools too. What's new is the layer above them. A general language model researched the biology of each target, chose the docking site, installed the programs itself from their public code repositories, combined them across 24 different workflows, and delivered a finished ranking without any human touching a single design decision. Since all the models used are open-source, the report argues, such campaigns are within reach for any lab.
The authors spell out the limits clearly themselves: There was no parallel campaign by human experts as a control, and they don't claim Claude's designs are better than what specialists would achieve with the same tools and budget. For four of the six contest targets, the contest results were in Claude's reading list.
The only thing measured was whether the designs bind, not their actual spatial shape and not their biological effect. Not a single design was structurally resolved, and all the binding models shown are predictions. Each combination of model, format, and target also ran exactly once, so model differences and chance can't be separated. And a large share of the human expertise sits in the prompt itself.
Anthropic has published these prompts, design data, and both measurement datasets on Hugging Face, so the campaign can in principle be reproduced and the dataset used as a benchmark.
Read on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: "AI Radar" — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI