摘要:在这篇文章中,我们分享了两项成果,展示 Claude 如何帮助生命科学家加快研究进度。在第一项成果中,我们测试了 Claude 从零设计蛋白质结合物的能力,这是代表药物设计流程早期阶段的关键任务,历来需要专家针对每个靶点花费数周或数月时间。Claude(Mythos Preview 和 Opus 4.8)针对 15 个靶点设计了蛋白质结合物,其中 14 个靶点取得成功。根据设置不同,其单个设计中有 22% 至 35% 成功结合,而当前蛋白质设计项目中的典型成功率仅为 10-15%。其一些最强设计在结合紧密程度上比此前发表的最佳结果高出数倍。在第二项成果中,我们评估了 Claude 能否加速化学分析。Claude Opus 5 是一款普遍可用的模型,我们向其提供了 NMR 和 LC-MS 数据(这些数据可让化学家评估所处理化合物的身份和纯度)。仅凭合同实验室的原始文件和一句两句话的提示词,Claude 分别在 23 分钟和 19 分钟内返回了完成结果,在氢原子计数和纯度方面与实验室自身的分析结果一致(96.4% 对 96.33%)。这些示例表明,Claude 如何减少当前在复杂科学任务上取得进展所需的时间和计算专业知识。AI 驱动的发现速度在过去几个月里有所加快。这些发现大多集中在验证相对较快的领域。例如,在数学领域,智能体已开始攻克未解问题:悬而未决数十年的 Erdős 问题正以每月数个的速度被解决,我们最近还分享了 Claude 如何改进了黎曼 zeta 函数上一个长期存在的下界。
AI 模型也开始加速推动实验科学领域的进展,在这些领域中,验证结果往往更加复杂且成本更高,例如生命科学领域。在本文中,我们分享了两项关于 Claude 科学能力的实验结果。首先,我们展示了针对 Claude 在蛋白质设计任务中表现的研究发现,结果表明 Claude 能够针对多种靶点设计蛋白质结合剂,其水平与顶尖人类专家相当(甚至更优)。其次,我们分享了 Claude Opus 5 在一项分析化学任务中的表现,展示了通用型模型如何支撑研究中常规且耗时的环节。
下文所述的蛋白质设计与分析化学任务,代表了药物开发早期阶段部分典型工作内容。加速这些阶段,是我们为端到端提速药物开发所付出的更大努力中的一个组成部分,而这一努力的许多方面,更多与政策和运营层面的瓶颈相关,而非核心科学能力的提升。
我们今天分享的结果,是通过 Mythos 与 Opus 模型的组合获得的。虽然生命科学研究任务目前在我们最强大的模型中被屏蔽,但我们的最高优先事项之一是为科学家推出访问计划,我们预计很快会分享更多相关信息。与此同时,Opus 5 仍是我们最强大的通用可用模型。
Claude 设计蛋白质
当我们发布 Claude Mythos 5 时,我们提到正在用该模型进行实验,以加速药物设计流程中的部分环节。作为这项工作的持续组成部分,我们一直在研究 Claude 为多个蛋白质靶点设计微型结合蛋白的能力。微型结合蛋白是一种小分子蛋白质,设计目的是紧密附着在靶蛋白上。结合是现代药物发挥作用的主要机制之一:药物附着到靶点上,对其进行抑制、激活或递送某种物质。设计一种新的结合蛋白(即从头设计)历来需要蛋白质工程师针对每个靶点花费数月时间进行计算、优化和筛选。近年来,能够设计蛋白质并对其结合可能性进行排序的机器学习模型大大加快了蛋白质设计流程。但这些模型通常仍需要计算专家耗费数天(往往甚至数周)进行繁琐的编排工作。此外,虽然像 Claude 这样的通用推理模型可以帮助专家和非专家更高效地进行蛋白质的计算设计,但在湿实验室(科学家在此对化学品、药物和其他生物物质进行实际测试)中验证这些数据仍然需要数周时间。
我们现在已经收到了这些实验中第一个实验的湿实验室数据,这是一项使用 Claude Opus 4.8 和 Mythos Preview 针对 15 个靶点开展的多臂蛋白质设计实验。我们的外部评估机构 Adaptyv Bio 和 Twist Bioscience 独立在实验室中生产并测试了 Claude 的设计,结果发现,在我们设计所针对的 15 个靶点中,Claude 成功为其中 14 个靶点设计出了结合蛋白。这些成果包括针对至少 6 个靶点的高亲和力结合蛋白,以及至少 4 个靶点的结合蛋白达到或超过了已报道的最佳亲和力。亲和力是衡量蛋白质与其靶点结合强度的指标;高亲和力结合蛋白通常需要达到治疗效果,因为它们能使药物在较低剂量下就发挥作用,从而降低副作用风险并减少制造成本。
Mythos Preview 和 Opus 4.8 在 48 小时会话中同时针对所有靶点进行设计时,整体命中率(即实际成为结合蛋白的设计占比)分别达到 26.7% 和 22.6%。目前蛋白质设计领域的典型命中率为 10% 至 15%。
在评估了 Claude 针对多靶点进行设计的能力后,我们希望了解让它一次专注于单一靶点是否能提升其表现,尤其是考虑到这更符合蛋白质工程师通常采用的工作方式。事实上,我们发现 Mythos Preview 在使用多个 24 小时会话分别针对每个靶点进行设计时,整体命中率达到 35.1%。

本次设计活动在极少人工介入的情况下完成,仅依靠我们在初始提示词中提供给 Claude 的信息。我们预计,在专业蛋白质设计师手中,这种方法将产生更出色的结果,尤其是当他们为 Claude 提供主动指导以及对中间结果的反馈时。
设计活动
我们首先选择了多个蛋白质设计基准中常用的靶点来启动设计活动,其中包括 Adaptyv Bio 的 BenchBB 全部靶点。由于这些靶点已被广泛研究,我们可以将结果与已发表的命中率和亲和力数据进行对比。我们还从 Adaptyv Bio 最近的竞赛中选取了两个新靶点——15-PGDH 和 GDF-8,以确保 Claude 能够在不依赖训练数据中已有成功案例或在线搜索的情况下针对靶点进行设计(对于所有靶点,我们均要求 Claude 检查并确保其设计具有原创性)。
随后,我们在 Claude Science 中提示 Claude 针对这些靶点设计蛋白质结合蛋白。为此,我们采用了两种方法。第一种是多靶点模式,即 Claude 在单个 Claude Science 会话中同时针对所有靶点进行设计。第二种是单靶点模式,即每个会话只针对一个靶点,且所有靶点的会话并行运行。
我们以多目标模式运行了 Opus 4.8 和 Mythos Preview,使用了 48 小时的墙钟时间和最多 12,500 个 NVIDIA H100 小时的计算资源来运行专门的蛋白质设计和折叠模型。我们还以单目标模式运行了 Mythos Preview,每个目标使用 24 小时的墙钟时间和最多 2,500 个 NVIDIA H100 小时的计算资源。5
为了模拟典型蛋白质设计活动中可用的资源,我们为 Claude 提供了以下内容:
- 一份详尽的蛋白质设计提示词6,该提示词也包含在智能体上下文中;
- 互联网访问权限以及一个资源库(例如关于蛋白质设计的论文);
- 用于 Google Drive、Slack、Gmail 和 BioRxiv 的连接器;
- 用于运行专门蛋白质设计和折叠模型的 GPU 访问权限;
- 在规定时间内不限制 token 和子智能体预算,并启用了快速模式。
在向 Claude 提供提示词后,我们让模型自主执行。在启动这些任务后,我们没有提供任何额外的科学、技术或操作指导。
我们唯一的参与是批准访问请求(例如网络访问请求)并监控基础设施以确保会话正常运行。Claude 完成了设计结合剂的所有工作,这些工作通常需要人类操作员花费数周时间。它自行选择在每个蛋白质靶标上的哪个位置进行设计;通过编排多个结构设计、序列设计和共折叠模型(这些模型可一次性预测蛋白质及其结合对象的结构)来生成候选结构和序列;通过多轮计算机模拟优化来运行这些设计;并进行计算筛选,以寻找新颖、多样化的候选分子,这些候选分子应能够表达、保持可溶性并实现结合。
对于 15 个靶标中的每一个,我们要求 Claude 设计 30 个蛋白质结合剂。Claude 通过操作该领域已经使用的公开可用的专业蛋白质设计和共折叠模型来完成这项工作。Claude 的设计随后被发送给 Adaptyv Bio 和 Twist Bioscience 进行验证。

Claude 在各靶点上的表现
到这项工作结束时,我们利用总计 1,320 个设计,针对 15 个靶点中的 14 个生成了 354 个结合蛋白。这是对公开可用的从头蛋白质设计语料库的重要贡献;例如,最大的两个数据集——proteinbase.com 和 Overath 等人整理的集合——包含针对 40 个靶点的 5,700 个设计中的约 770 个结合蛋白。下面,我们分享三个突出 Claude 能力的示例,以及一个展示其局限性的示例。更多细节请参阅我们的技术报告(链接)。
Claude 的设计与 Adaptyv Bio 蛋白质设计竞赛中的参赛作品具有竞争力
我们发现,对于 Adaptyv Bio 已举办过竞赛的靶点,Claude 在命中率和亲和力两个指标上的表现达到或超过了顶尖参赛者的水平。针对 RBX1(一种驱动特定调控蛋白靶向降解的小蛋白),Mythos Preview 在单靶点模式下实现了 40% 的命中率,而参赛者的命中率为 3.7%。其排名最高的设计是一个高亲和力结合蛋白,表现优于获胜设计——后者出自 245 个参赛设计之中。

Claude 针对 TNFα(一个具有挑战性且与治疗相关的靶点)设计了物种交叉反应性结合蛋白。
有趣的是,成功攻克 TNFα 这一靶点的是 Opus 4.8,而非 Mythos Preview。TNFα 是免疫系统释放的一种信号蛋白,用于触发炎症反应,阻断它是包括 Humira 在内的一些最具影响力药物的治疗基础。由于 TNFα 具有多聚体结构,需要靶向由两个蛋白形成的沟槽中的结合位点,因此这是一个颇具挑战性的设计靶点。尽管 Mythos Preview 未能成功,Opus 4.8 却设计出了多种结合蛋白,其中一些具有跨物种活性,可结合人、食蟹猴和小鼠的 TNFα,这对于开展动物研究至关重要。我们尚不清楚为何 Opus 4.8 在该靶点上成功而 Mythos Preview 却失败。我们在评估模型能力时是进行整体评估的。鉴于蛋白质设计固有的复杂性,在特定领域出现一个整体能力较弱的模型反而胜过整体能力更强的模型的情况,并不令人意外。

Claude 设计了具有 β-折叠的折叠多样性结合蛋白
大多数计算设计的结合蛋白是 α-螺旋束,这是一种由卷曲结构组成的蛋白质二级结构。β-折叠则要求延伸的氨基酸链并排排列,设计难度更大,且更容易出现错误折叠和聚集(即蛋白质分子相互粘连,而非保持分离并正确折叠的状态)。Claude 针对六个靶点设计了 15 个经确认的结合蛋白,其中 β-链含量至少达到 20%,这展示了其对蛋白质结构进行推理的能力。

Claude 在某些靶点上遇到了困难
某些靶点对 Claude 而言仍是挑战,包括 BBF-14 和麦芽糖结合蛋白(MBP)。BBF-14 是一种自然界中不存在的 β-桶状蛋白:它本身就是从头设计的产物,正因其新颖性,如今被用作结合子设计的基准。MBP 的结构也格外困难。MBP 是一种大型、柔性的细菌蛋白,表面光滑且亲水,使其成为良好的实验室试剂。这导致结合子几乎找不到可抓取的位点。尽管如此,Claude 仍成功生成了三个独立的 BBF-14 结合子——每个设计分支各一个,且各自基于不同的骨架——亲和力为中等水平(亚微摩尔至微摩尔级)。然而,针对 MBP,90 个设计中没有一个被确认与靶点结合,尽管其中一个表现出微弱但可重复的结合信号。

为了更好地了解 Claude 在这些设计活动中的表现,我们计划在实验之后进行更全面的表征,以确认我们的命中率和亲和力测量结果。与此同时,我们正在分享这些活动中使用的提示词,以及我们生成的所有体外和计算机模拟数据。7
智能体驱动的生物发现具有双重用途
AI 模型日益自主的研究能力所带来的提升,无疑将加速人类疗法的开发和基础科学发现。然而,这类能力也具有双重用途:如果没有稳健的安全措施,它们可能使恶意行为者得以开展危险研究,例如开发生物武器。在我们通过可信访问计划安全地交付这些能力的同时,蛋白设计及其他双重用途的研究生物学能力仍不对 Claude Fable 5 的普通访问开放。不过,正如你将在下文看到的,我们的 Opus 级模型能够完成卓越的科学工作。
Claude 运行分析化学工作流程
在蛋白质结合剂实验中测试了 Claude 设计新分子的能力,而第二个实验则测试了它解读已有分子测量结果的能力。表征一种化合物是繁琐的工作;与蛋白质设计非常相似,它要求化学家进行多轮测量、分析和迭代。例如,每次化学家创造出一个分子,他们必须确定该分子是否是他们预期要生产的产物,以及其纯度如何。这通常通过核磁共振(NMR)波谱法来完成。NMR 谱图是一系列峰,每个峰对应分子中某个位置的氢原子,或一组等价的氢原子。峰的位置告诉化学家每个氢原子连接在什么基团上,峰的大小则表示它代表多少个氢原子。确认结构是合成化学中最耗时的步骤之一;对于每一种化合物,化学家都必须手动将谱图中的每个峰与所提出结构中的某个原子进行匹配。
另一种主要用于评估纯度的技术是液相色谱-质谱联用(LC-MS),它首先将样品通过色谱柱分离成各个组分,然后根据其紫外吸收记录各组分的含量,最后测量每个组分的分子质量。对于这两种技术,仪器运行本身只需几分钟(常规质子 NMR 谱图约需两到三分钟;一次 LC-MS 运行约需 10 分钟)。繁琐的部分在于分析输出结果。在 NMR 和 LC-MS 运行之后,每台仪器都会生成一个制造商专有格式的原始文件,该文件旨在用该制造商(或其他专业)的软件打开。8
鉴于处理和解读这些文件极为繁琐,我们想看看像 Claude Opus 5 这样的通用模型在此类任务上表现如何。9 仅凭合同实验室提供的常规质控样品原始文件,以及一段简短的自然语言提示词,10 在无供应商软件、无操作员介入的情况下,Claude 在 Claude Science 环境中并行工作,分别在 23 分钟和 19 分钟内返回了处理后的 NMR 和 LC-MS 结果。其结果与实验室自身的处理结果一致——每个峰的氢计数与实验室结果相差在 0.08 ¹H 以内,其测得的纯度为 96.4%,而实验室结果为 96.33%。

对于 NMR 数据,Claude 将仪器输出的原始数据转换为校准谱图和一张包含 18 个峰的表格,每个峰都有对应的氢计数。接下来,它像化学家一样,将四个宽峰标记为可能连接在氮或氧上的氢。随后它提出了标准验证方法:向 NMR 样品中加入重水,这会置换掉这些氢,使其峰缩小或消失。(实验室在首次测量三天后也独立进行了同样的验证。)拿到重水实验的原始文件后,Claude 量化了数据中的变化,发现并纠正了其首次解读中的一处夸大(第一遍它报告所有四个标记峰均已消失,但其自身复核显示只有两个消失),并最终得出了与实验室操作员相同的结论。

LC-MS 仪器文件采用一种未公开的供应商专有格式。Claude 先自行弄清了数据的编码方式,然后在分析任何内容之前,通过复现仪器自身记录的全部 2,664 次扫描的总数,确认了自己正确读取了文件。随后,它输出了化学家所期望的所有结果:分离图谱、质谱和紫外光谱、纯度表、化合物的分子量,以及可用于读取此类文件的可复用代码,并附上了它自己关于结果可信度的注意事项清单(例如,它指出这类仪器给出的质量数只精确到最接近的整数单位)。
通常情况下,化学家需要手动完成所有这些分析。对于治疗相关的小分子,每个样品通常需要半小时到一小时(实验室自己的记录显示,每张 NMR 谱图大约需要两分钟的人工处理,LC-MS 报告则在样品加载到仪器后约两小时出具)。Claude Science 在这 25 分钟内并行处理并解读了这两个文件。Claude 还在同一时间内生成了一份书面报告,而该实验室针对此样品的最终报告是在采集到第一张谱图四天后才完成的——考虑到他们一次只分析一个分子,且期间可能穿插其他工作,这是一个相当标准的延迟时间。
除了效率提升之外,Claude 的运行还展现出一定程度的科学判断力,例如它提出了与外包实验室独立开展的后续实验完全相同的建议。随着这些模型不断改进,我们预计 Claude 的科学判断力将变得更加敏锐。
如果你想在 Claude Science 中亲自尝试,可以给 Claude 一个原始的 NMR 或 LC-MS 文件,并让模型确认该化合物的身份和纯度。
结论
上述两个例子都展示了 AI 模型如何通过减少科学发现所涉及的专业知识、成本和时间来加速生命科学研究。在化学领域,Claude 正在将化学家历来手动完成的分析工作自动化。在蛋白质设计领域,Claude 能够以极少的输入端到端地执行结合剂设计任务,产出的结合剂可与之前发表的最佳设计相媲美甚至更胜一筹。
蛋白质微型结合剂并非药物的标准治疗模式,即便是单克隆抗体和小分子这类常见药物模式,设计高亲和力结合剂也只是生成类药分子的第一步。不过,我们认为这项工作具有奠基性意义,并正在对其进行扩展,以便 Claude 能够端到端地运行涵盖所有药物模式的完整开发流程。
延伸阅读
以下文档列表提供了关于上述结果的更多技术深度和更详细信息:
- 蛋白质设计活动的提示词和数据;
- 蛋白质设计技术报告;
- 化学分析技术报告。
脚注
- 我们认为,如果结合剂的平衡解离常数达到个位数纳摩尔或更低(KD < 10 nM),则视为高亲和力结合剂。
- 通过计算 https://proteinbase.com/ 上从头设计的蛋白质结合剂总体命中率得出。
- 这包括批准 Claude 提出的某些请求(例如网络访问请求、代码执行请求)、解决蛋白质设计会话之外的基础设施问题,以及对生成的候选设计进行排序以供实验验证。
- 我们总共选择了 16 个靶点(相关情况下默认物种为人类)。按字母顺序排列,它们分别是:15-PGDH、BBF-14、BHRF1、Cas9、EGFR、GDF-8(潜伏态)、GDF-8(成熟态)、IL-7Rα、麦芽糖结合蛋白(MBP)、尼帕病毒糖蛋白 G、PD-L1、RBX1、TNFα、TREM2、TrkA 和 VEGF-A。我们展示了 16 个靶点中 15 个的结果,因为其中一个靶点 GDF-8(成熟态)的实验数据因靶点聚集和非特异性粘附而无法得出明确结论。
- 15-PGDH 和潜伏态 GDF-8 未用于多靶点模式。Opus 4.8 以单靶点模式针对三个靶点运行:TNFα、潜伏态 GDF-8 和成熟态 GDF-8。
- 包含约 30,000 个 token,在此共享。
- 提示词、所设计蛋白质复合物的计算模型以及实验数据均可在此找到。
- 对于 NMR,指波谱仪记录的原始信号(“自由感应衰减”);对于 LC-MS,指仪器的二进制运行文件。两者均为仪器的原始专有文件。
- 我们此前曾分享过 Claude 在分析 NMR 数据方面、与标准软件对比的性能表现工作。
- NMR 提示词全文如下:“我有一段原始 1H FID 数据:请处理它:进行傅里叶变换、相位校正、基线校正。显示谱图。然后挑选峰并积分:给我一个包含 δ(ppm)、多重性、J(Hz)和积分值的表格。”LC-MS 提示词为:“处理原始 LCMS 文件:提取色谱图和质谱图,并用图表进行总结。”
Summary: In this post, we share two results that show how Claude can help life scientists increase the pace of their research. In the first, we tested Claude’s ability to design protein binders from scratch, a key task representative of the early parts of the drug design process and one that has historically taken a specialist weeks or months per target. Claude (Mythos Preview and Opus 4.8) designed protein binders against 15 targets, and succeeded against 14 of them. Between 22% and 35% of its individual designs bound successfully, depending on the setup, compared to the 10-15% that is typical in protein design campaigns today. Some of its strongest designs bound several times more tightly than the best previously published result. In the second example, we evaluated whether Claude can accelerate chemical analysis. Claude Opus 5, a generally available model, was given NMR and LC-MS data (the data that allows chemists to assess the identity and purity of the compounds they work with). Provided with only a contract lab’s raw files and a two-sentence prompt, Claude returned finished results in 23 and 19 minutes, matching the lab’s own analysis on hydrogen counts and purity (96.4% versus 96.33%). These examples demonstrate how Claude can reduce the time and computational expertise currently required to make progress on complex scientific tasks.
The pace of AI-enabled discoveries has quickened over the past few months. The bulk of these discoveries have been in areas where verification is relatively fast. In mathematics, for example, agents have begun to work their way through unsolved problems: Erdős problems that have stood for decades are falling at a rate of several a month, and we recently shared how Claude improved on a longstanding lower bound on the Riemann zeta function.
AI models are also beginning to hasten progress in experimental fields where verifying the results is more complex and expensive, such as in the life sciences. In this post, we share the results of two experiments into Claude’s scientific capabilities. First, we present findings from our investigation into Claude’s performance on a protein design campaign, showing that Claude can design protein binders against a variety of targets as well as (or even better than) leading human experts. Second, we share how Claude Opus 5 performed on an analytical chemistry task, demonstrating how general-access models can support the routine and time-intensive aspects of research.
The protein design and analytical chemistry tasks described below are representative of the work that makes up some parts of the early stages of the drug development process. Accelerating these phases is one component of our much larger effort to speed up drug development end-to-end, many aspects of which have more to do with policy and operational bottlenecks than with improvements in core scientific capabilities.
The results that we’re sharing today were obtained with a combination of our Mythos and Opus models. While life science research tasks are currently blocked in our most capable model, one of our highest priorities is to launch an access program for scientists, and we expect to share more on this soon. In the meantime, Opus 5 remains our most capable generally available model.
Claude designs proteins
When we announced Claude Mythos 5, we shared that we were experimenting with the model to accelerate parts of the drug design process. As an ongoing part of this work, we have been investigating Claude’s ability to design minibinders for multiple protein targets. A minibinder is a small protein designed to latch tightly onto a target protein. Binding is how a large proportion of modern medicines work: they attach to a target and inhibit, activate, or deliver something to it. Designing a new binder (known as de novo design) has historically taken protein engineers months of computation, optimization, and screening per target.
In recent years, machine-learning models that can design proteins and rank which are most likely to bind have greatly expedited the protein design process. But these models still generally require days (and often weeks) of laborious orchestration by computational experts. And although general reasoning models like Claude can help both experts and non-experts more efficiently design proteins computationally, validating that data in a wet lab (where scientists physically test chemicals, drugs, and other biological substances) still takes weeks.
We have now received wet lab data back for the first of these experiments, a multi-arm protein design campaign against 15 targets using Claude Opus 4.8 and Mythos Preview. Our external evaluators, Adaptyv Bio and Twist Bioscience, independently produced and tested Claude’s designs in the lab, finding that of the 15 targets we designed against, Claude successfully designed binders against 14 of them. These include high-affinity binders1 against at least six targets, and binders matching or exceeding the best reported affinity against at least four targets. Affinity is a measure of how strongly a protein binds to its target; high-affinity binders are generally needed to achieve a therapeutic effect because they make the drug effective at lower doses, reducing the risk of side effects and the cost to manufacture them.
Mythos Preview and Opus 4.8 achieve overall hit rates—how many of the designs are, in fact, binders—of 26.7% and 22.6%, respectively, when designing against all targets simultaneously in a 48-hour session. 10 to 15% is typical in protein design campaigns today.2
After assessing Claude’s ability to design against multiple targets, we wanted to understand whether having it focus on a single target at a time would improve its performance, especially given that this better represents the approach typically taken by a protein engineer. Indeed, we found that Mythos Preview achieves an overall hit rate of 35.1% when designing against each target separately using multiple 24-hour sessions.

This campaign was carried out with minimal human involvement3 beyond the information we provided Claude in our initial prompt (link). We expect that in the hands of expert protein designers this approach would yield even stronger results, especially if they give Claude active guidance and feedback on intermediate results.
The campaign
We began our protein design campaign by selecting multiple targets4 that are commonly used in protein design benchmarks, including all of Adaptyv Bio’s BenchBB. Because these targets have been studied extensively, we can compare our results against published hit rates and affinities. We also chose two novel targets, 15-PGDH and GDF-8, from Adaptyv Bio’s most recent competitions to ensure Claude was able to design against targets without drawing upon pre-recorded successes in its training data or from online search (for all targets, we required Claude to check for and ensure that its designs were original).
We then prompted Claude to design protein binders against these targets in Claude Science. For this, we took two approaches. The first was a multi-target mode, where Claude designed against all targets simultaneously in a single Claude Science session. The second was a single-target mode, in which each session addressed one target and sessions for all targets ran in parallel.
We ran Opus 4.8 and Mythos Preview in multi-target mode with 48 hours of wall time and up to 12,500 NVIDIA H100 hours of compute for running specialized protein design and folding models. We also ran Mythos Preview in single-target mode with 24 hours of wall time and up to 2,500 NVIDIA H100 hours of compute for each target.5
To emulate the resources available during a typical protein design campaign, we gave Claude the following:
- An extensive protein design prompt6 that was also included in the agent context;
- Access to the internet and a corpus of resources, such as papers, on protein design;
- Connectors for Google Drive, Slack, Gmail, and BioRxiv;
- Access to GPUs for running specialized protein design and folding models;
- No limits on token and sub-agent budget within the allotted time, and fast mode enabled.
After giving Claude the prompt, we left the model to execute autonomously. We provided no additional scientific, technical, or operational guidance after we initiated the campaigns.
Our only involvement was granting access approvals (such as network access requests) and monitoring the infrastructure to ensure the sessions were running. Claude conducted all of the work that goes into designing a binder, which can take a human operator weeks. It chose where on each protein target to design against; generated candidate structures and sequences by orchestrating several structure design, sequence design, and co-folding models (models that predict the structure of a protein, together with whatever it binds, in a single pass); ran the designs through multiple cycles of in silico optimization; and computationally screened for novel, diverse candidates that would express, stay soluble, and bind.
For each of the 15 targets, we asked Claude to design 30 protein binders. Claude did this by operating publicly available specialist protein design and co-folding models that the field already uses. Claude’s designs were then sent to Adaptyv Bio and Twist Bioscience to validate.

Claude’s performance on the targets
By the end of this effort, we produced 354 binders against 14 of 15 targets using a total of 1,320 designs. This represents a significant contribution to the total corpus of publicly available de novo protein designs; for example, the two largest collections, proteinbase.com and the collection curated by Overath et al., consist of approximately 770 binders out of 5,700 designs against 40 targets. Below, we share three examples highlighting Claude’s capabilities, and one showing its limitations. You can find more detail in our technical report (link).
Claude’s designs are competitive with entries in Adaptyv Bio’s protein design competition
We found that for the targets Adaptyv Bio has run competitions for, Claude performs at or beyond the level of the top participants on both hit rate and affinity. Against RBX1 (a small protein that drives the targeted destruction of specific regulatory proteins), Mythos Preview in single-target mode achieved a 40% hit rate, compared to a 3.7% hit rate among participants. Its top-ranked design was a high-affinity binder that outperformed the winning design, which was among 245 designs entered.

Claude designs species cross-reactive binders against TNFα, a challenging, therapeutically relevant target
Interestingly, Opus 4.8, and not Mythos Preview, succeeds on TNFα, a target multiple expert groups have struggled with. TNFα is a signaling protein released by the immune system to trigger inflammation, and blocking it is the therapeutic basis for some of the most impactful drugs ever made, including Humira. It’s a challenging target to design against because of its multimeric structure, which requires targeting a binding site in the groove formed by two proteins. Although Mythos Preview was unsuccessful, Opus 4.8 designed multiple binders, including some that worked across species, binding human, cynomolgus monkey, and mouse TNFα, which is important for conducting animal studies. We’re not sure why Opus 4.8 was successful on this target and Mythos Preview was not. When we assess our models capabilities, we do so holistically. Given the inherent complexity of protein design, it’s unsurprising that there would be specific areas where an overall less capable model could still outperform one that was generally more capable.

Claude designs fold-diverse binders with β-sheets
Most computationally designed binders are bundles of α-helices, a protein secondary structure consisting of coils. β-sheets, in which extended strands of amino acids must line up side by side, are harder to design and more prone to misfolding and aggregation (when protein molecules stick to each other instead of staying separate and properly folded). Claude designed 15 confirmed binders across six targets that contain at least 20% β-strand, demonstrating its ability to reason about protein structure.

Claude struggled against some targets
Certain targets remained a challenge for Claude, including BBF-14 and maltose-binding protein (MBP). BBF-14 is a β-barrel-shaped protein that does not exist in nature: it was itself de novo designed, and it is now used as a benchmark for binder design precisely because of its novelty. MBP’s structure is also especially difficult. MBP is a large, flexible bacterial protein with a smooth, water-loving surface that makes it a good lab reagent. This leaves a binder very little to grab on to. Claude still managed to produce three independent BBF-14 binders—one from each design arm, and each built on a different backbone—with modest (sub-micromolar to micromolar) affinities. Against MBP, however, none of the 90 designs was confirmed to have bound to the target, although one demonstrated a weak, reproducible binding signal.

To better understand how well Claude performed across these design campaigns, we intend to follow our experiments with more extensive characterization to confirm our hit rates and affinity measurements. In the meantime, we are sharing the prompts we used for these campaigns, as well as all in vitro and in silico data we generated.7
Agentic biological discovery is dual-use
The uplift provided by the increasingly autonomous research capabilities of AI models will undoubtedly speed the development of human therapies and fundamental scientific discoveries. However, such capabilities are also dual-use: without robust safety measures, they could enable bad actors to perform dangerous research, such as the development of bioweapons. As we work to deliver these capabilities safely via trusted access programs, protein design and other dual-use research biology capabilities remain unavailable for general access in Claude Fable 5. However, as you’ll see below, our Opus class models are capable of remarkable scientific work.
Claude runs the analytical chemistry workflow
Where the protein binder campaign tested Claude’s ability to design new molecules, the second experiment tested its ability to interpret measurements of molecules already made. Characterizing a compound is cumbersome work; much like protein design, it requires chemists to perform many rounds of measurement, analysis, and iteration. For example, every time a chemist creates a molecule, they must establish whether it is what they intended to produce and how pure it is. This is typically done with nuclear magnetic resonance (NMR) spectroscopy. An NMR spectrum is a series of peaks, each corresponding to a hydrogen atom, or a group of equivalent hydrogens, somewhere in the molecule. The location of the peaks shows chemists what each hydrogen atom is attached to, and the size of the peak shows how many hydrogens it represents. Confirming a structure is one of the most time-consuming steps in synthetic chemistry; for every compound, a chemist has to match each peak in the spectrum to an atom in the proposed structure by hand.
The other technique, used mainly to assess purity, is liquid chromatography–mass spectrometry (LC-MS), which first separates the sample into its individual components as they flow through a column, then records how much of each is present based on its ultraviolet absorbance, before measuring the molecular mass of each one. For both techniques, the instrument run itself takes only a few minutes (two to three for a routine proton NMR spectrum; about 10 for an LC-MS run). The tedious part is analyzing the output. After NMR and LC-MS are run, each instrument produces a raw file in the manufacturer’s own format that is meant to be opened in that manufacturer’s (or other specialist) software.8
Given how painstaking it is to process and interpret these files, we wanted to see how a generally available model such as Claude Opus 5 would perform at this task.9 Supplied with only a contract lab’s raw files for a routine quality-control sample and a short plain-language prompt,10 with no vendor software and no operator, Claude, working within Claude Science, returned processed NMR and LC-MS results in 23 and 19 minutes, respectively, working in parallel. Its results matched the lab’s own processing—hydrogen counts per peak were within 0.08 ¹H of the lab’s, and its purity was measured at 96.4% versus the 96.33% of the lab.

For the NMR data, Claude converted the raw data from the instrument into a calibrated spectrum and a table of 18 peaks, with a hydrogen count for each. Next, as a chemist would, it flagged four broad peaks as hydrogens that were probably attached to nitrogen or oxygen. It then proposed the standard check: add heavy water to the NMR sample, which swaps those hydrogens out so their peaks shrink or vanish. (Independently, the lab had run this same check three days after the first measurement.) Given the raw file from the heavy-water run, Claude quantified what had changed in the data, caught and corrected an overstatement in its first reading (its first pass reported that all four flagged peaks had disappeared, but its own self-check showed that only two had), and arrived at the same conclusion as the lab’s operator.

The LC-MS instrument files use an undocumented vendor format. Claude worked out how the data was encoded, then confirmed it had read the file correctly by reproducing the instrument's own recorded totals for all 2,664 scans before analyzing anything. It then delivered all the outputs a chemist would expect: the separation trace, mass and UV spectra, a purity table, the compound’s molecular mass, as well as reusable code for reading such files, alongside its own list of caveats about the trustworthiness of the results (it noted, for example, that this class of instrument gives the mass only to the nearest whole unit).
Ordinarily, a chemist does all this analysis by hand. This typically takes half an hour to an hour per sample for a therapeutically relevant small molecule (the lab’s own records show about two minutes of hands-on processing per NMR spectrum, with the LC-MS report following about two hours after the sample was loaded onto the instrument). Claude Science processed and interpreted both files in parallel within those 25 minutes. Claude also produced a written report in that time, whereas the lab’s finished report for this sample arrived four days after the first spectrum was acquired—a fairly standard lag time given that they analyze molecules one at a time, and work may crop up in between.
Beyond the increased efficiency, Claude’s run also showed a degree of scientific judgment, for instance in proposing the very same follow-up experiment the contract lab had independently run. As these models continue to improve, we expect Claude’s scientific judgment to become more acute.
To try this yourself in Claude Science, give Claude a raw NMR or LC-MS file and ask the model to confirm the compound’s identity and purity.
Conclusion
Both of the examples above demonstrate how AI models can accelerate research in the life sciences by reducing the expertise, cost, and time involved in scientific discovery. In chemistry, Claude is automating analyses that chemists have historically done by hand. In protein design, Claude can execute binder design campaigns end-to-end with minimal input, producing binders that match or surpass the best previously published designs.
Protein minibinders are not a standard therapeutic modality for drugs and even for the common drug modalities, such as monoclonal antibodies and small molecules, designing a high-affinity binder is just the first step in the process of generating a drug-like molecule. However, we view this work as foundational, and are extending it so that Claude can run the entire development process end-to-end across all drug modalities.
Further reading
Below is a list of documents that provide further technical depth and more detailed information about the results described above:
- Prompts and data for the protein design campaign;
- Protein design technical report;
- Chemical analysis technical report.
Footnotes
- We consider binders to be high-affinity if they have at most single-digit nanomolar equilibrium dissociation constants (KD < 10 nM).
- Derived by calculating overall de novo protein binder hit rates on https://proteinbase.com/.
- This consisted of approving certain requests made by Claude (e.g., network access requests, code execution requests), resolving infrastructure issues outside of the protein design sessions, and ordering the generated designs for experimental validation.
- In total, we selected 16 targets (the default species was human where relevant). In alphabetical order, they are: 15-PGDH, BBF-14, BHRF1, Cas9, EGFR, GDF-8 (Latent), GDF-8 (Mature), IL-7Rα, Maltose Binding Protein (MBP), Nipah virus Glycoprotein G, PD-L1, RBX1, TNFα, TREM2, TrkA, and VEGF-A. We present results for 15 of 16 targets, as the experimental data for one target, GDF-8 (Mature), were inconclusive due to target aggregation and non-specific stickiness.
- 15-PGDH and latent GDF-8 were not used in multi-target mode. Opus 4.8 was run in single-target mode against three targets: TNFα, latent GDF-8, and mature GDF-8.
- Consisting of roughly 30,000 tokens, shared here.
- Prompts, computational models of the designed protein complexes, and experimental data can all be found here.
- For NMR, the raw signal (a “free induction decay”) as written by the spectrometer; for LC-MS, the instrument’s binary run file. Both are raw proprietary instrument files.
- We have previously shared work on Claude’s performance analyzing NMR data against standard software.
- The NMR prompt, in full: “i have a raw 1H FID: process it: FT, phase, baseline-correct. show me the spectrum. then pick peaks and integrate: give me a table with δ (ppm), multiplicity, J (Hz), and integral.” The LC-MS prompt: “Process the raw LCMS file: extract chromatograms and mass spectra, and summarize with figures.”