我们正与世界一流的合成化学家、计算化学家和分析化学家合作,让 Claude 在化学领域表现更出色。在这篇文章中,我们将分享这项工作的初步成果——Anthropic 化学家 David Kamber 研究了 Claude 如何处理化学家最常用的分析输入:核磁共振波谱。在处理分子时,化学家需要在白板上的手绘结构、仪器读数、数据库查询字符串以及专利和出版物中的技术符号之间切换。每种表示方式都编码了相同的底层化学信息,但每种都需要不同的熟练度。例如,咖啡因的草图能让化学家发现它与腺苷(人体困倦信号)的相似性,并预测它通过阻断受体来让我们保持清醒。然而,同样的草图却无法帮助化学家将其与其他外观几乎相同的分子区分开来。
理解化学家正在处理的是哪种分子至关重要。化学支撑着一切,从我们摄入的食物和药物,到我们的乳液、油漆和塑料。在相同的原子之间重新排列几个化学键,葡萄糖就会变成果糖——这两种分子共享相同的化学式,但通过完全不同的代谢途径被处理。将一个分子翻转成它的镜像,镇静剂就会变成致畸剂,正如沙利度胺灾难中所发生的那样。化学家的日常工作依赖于在适合特定任务的任何表示方式中正确解读这些信号。
在这些表示方式之间进行转换(从图中找出结构、将仪器读数与预期产物进行比对、以正确的符号查询数据库)既耗时,又无法大规模持续跟上——CAS(最大的化学物质注册库)收录了超过 2.9 亿种已公开的物质,并且每天新增约 15,000 种。
人工智能完全有能力承担这一研究负担,但在化学领域,这很大程度上仍停留在理想层面。多年来,机器学习工具一直被视为逆合成分析(即从目标分子逆向推导出更简单的前体,以规划其合成路径)、反应预测和性质评估等领域的变革性力量,但这些工具所需的数据却难以获取——数据在无效结果方面稀疏、格式不统一,并且被锁定在订阅制期刊的付费墙之后(以及非结构化的支撑信息中)。逆合成分析就是一个典型例子——功能强大的 AI 工具已存在多年,但应用并不均衡,普通学术机构或小型实验室的化学家仍然很少使用它们。
即便如此,人工智能的进步终于开始触及化学领域。当今的前沿模型是多模态的,并且具备显式推理能力。它们可以直接从期刊图表或手绘草图中读取化学结构,而无需依赖预先整理的分子数据库。它们还能以实际发表的形式,读取方法部分或支撑信息中的实验细节。此外,它们可以逐步展示推理过程,这意味着化学家能够审核其输出结果。这些进展都无法消除该领域多年来一直描述的数据问题,但它们改变了在数据问题存在的情况下哪些问题变得可解。归根结底,我们的主张是适度的:Claude 已开始有意义地协助化学家完成日常的翻译、回忆和整合工作,这些工作是对他们专业判断的补充,并且我们计划持续扩展其辅助能力。今天,我们发布了第一份白皮书,旨在加速这项工作。它处理的是化学家最常用的分析输入:核磁共振波谱。
Claude 与 ChemDraw 在核磁共振波谱预测和结构解析方面的对比
完整版本可在此处查阅
几乎所有小分子——药物、农药、染料、香料、聚合物、DNA或蛋白质亚基,以及功能性无机或固态材料——之所以存在,都是因为化学家确定了它们的结构。由于这些分子无法用显微镜观察,化学家必须依赖光谱分析,用光、无线电波或磁场来探测分子。给定分子吸收、发射或偏转这种能量的方式,会为化学家提供一种模式,即光谱,他们可以借此阐明其结构。核磁共振波谱学——化学家为此依赖的经典技术之一——是合成化学中最耗时的步骤之一;对于每一种化合物,化学家都必须手动将光谱中的每个峰与所提出结构中的某个原子进行匹配。在本白皮书中,我们测试了 Claude 与当今化学家依赖的专用 NMR 软件相比表现如何。我们在 20 种化合物上测量了三个 Claude 模型(Opus 4.7、Opus 4.6、Sonnet 4.6),与 ChemDraw 和 MestReNova 进行对比,这些化合物选自模型训练截止日期之后发表的合成化学预印本,以避免选择偏差。ChemDraw 和 MestReNova 都进行正向预测,利用绘制的结构来模拟将产生的 NMR 谱图。除了正向预测,我们还希望看看 Claude 能否反向操作——从实验谱图出发,提出其背后的结构。这是更困难的任务,也是现有软件目前留给化学家去完成的工作。为建立评估,我们从模型训练截止日期后发布的 ChemRxiv 预印本中选取了 20 种化合物,从每篇论文中提取了首批完全表征的新分子。这 20 种化合物涵盖四个结构家族,每个家族五种化合物,选择每个家族是因为它涉及不同类型的 NMR 挑战。每个工具都获得了以 SMILES 字符串(化学家用来将分子输入软件的行文本表示法)编码的结构,并被要求预测每个氢和碳峰在 1D NMR 谱图(以 ppm 为单位测量化学位移的水平轴)上的落点。由于 NMR 样品溶解在液体中,并且溶剂(氯仿、DMSO 等)的选择会略微移动峰的位置,因此每个工具都被要求根据化学家在已发表论文中使用的溶剂来预测谱图。

由于语言模型的输出在不同运行间存在差异,每个 Claude 模型对每种化合物均查询三次并取平均值;ChemDraw 和 MestReNova 每次返回相同结果,因此各运行一次。随后我们将每个预测峰与其实验对应峰配对,并测量 ppm 差值。这些差值均落在化学家认为正确的范围内——氢谱为 ±0.20 ppm,碳谱为 ±1.0 ppm。

在氢谱方面,Opus 4.7 最为准确,平均误差为 ±0.079 ppm——远低于容差窗口的一半——且落在该窗口内的峰占比最高。在碳谱方面,Opus 4.7 与 MestReNova 基本持平,误差分别为 ±1.37 和 ±1.48 ppm;其余工具在两种元素上保持相同的排名顺序。Opus 4.6 表现中规中矩,而 Sonnet 4.6 最弱。两者之间的差距在一个公认极难测定的氢原子上最为明显——即氯代吡哒嗪家族中的 NH 质子,其真实位置落在 6.8 至 7.9 ppm 的狭窄区间内。Opus 4.7 的预测值略低但始终一致;Opus 4.6 的猜测分散在数个 ppm 范围内;Sonnet 4.6 则将其置于 10–13 区间,远高于其实际出现位置。

虽然 Opus 4.7 的表现与 ChemDraw 和 MestReNova 相当,但在预测氢的 NMR 峰形状以及峰间距方面,差距更大——这些特征同样包含化学家除峰位外会读取的结构信息。Opus 4.7 匹配实验报告裂分模式的频率高于其他任何工具,并且所有三个 Claude 模型在约 80% 的情况下能将子峰间距预测精确到半赫兹以内——而 ChemDraw 和 MestReNova 的这一比例仅为 26% 到 35%。Opus 4.7 在三次重复运行中也最为稳定:其平均误差在不同运行之间的波动幅度,甚至小于它与次优工具之间的差距。在此基础上,我们评估了反向预测(结构解析):能否从谱图中确定分子的结构?我们向 Opus 4.7 提出了 15 个解析问题,每个问题重复三次,要求它提出最多三个按排序的候选结构。每次提供该化合物的精确分子式(来自高分辨质谱)及其氢谱和碳谱 NMR。这 15 个问题按难度划分。八个较简单的目标——单环或双片段分子——仅提供分子式和谱图。七个较复杂的目标——稠环、螺环等——附带一个额外提示:参与反应的起始原料的结构。

Opus 4.7 仅凭光谱和分子式,在每次尝试中都成功恢复了全部八个较简单的结构。对于七个较难的目标,在给出起始原料提示的情况下,它在其中四个目标的三次运行中均返回了正确结构,在其余目标的三次运行中有两次返回了正确结构。最终,我们发现,对于常规数据预测,Opus 4.7——一个未经化学领域特定微调的通用模型——现在平均而言已与 ChemDraw 和 MestReNova 相当或更优。此外,Claude 还能反向工作,仅凭核磁共振数据提出结构。专用的结构解析软件已存在数十年,但通常需要二维核磁共振(一种具有两个轴的光谱,输出是等高线图而非一排峰)、专业培训和授权工具。而 Claude 仅凭化学家粘贴到聊天窗口中的相同高分辨质谱和一维峰列表即可完成,无需任何设置。
局限性
这项评估表明,通用模型可以与核磁共振软件竞争,甚至能使一维反向解析变得可行。但仍存在一些值得注意的局限性。
- 首先,评估规模较小——正向任务涉及四个骨架的 20 种化合物,反向任务涉及 15 种化合物——且每个骨架仅贡献单一类别的失败模式。因此,模型性能应被视为指示性而非精确性。
- 其次,对于最密集的反向目标,如果没有起始原料作为额外输入,模型可能会在推理过程中循环,而无法确定最终结构;这就是为什么七个较难的问题在提出时附带了起始原料结构,而非仅凭光谱。
- 第三,某些化学骨架未经测试。例如,慢交换 NH 杂芳烃(其 N–H 与溶剂交换速度足够慢,从而在核磁共振中留下尖锐峰的芳香环)仅通过氯哒嗪进行了采样,排除了相关体系(羟基吡啶、氨基噻唑以及其他 DMSO-d₆ 中具有 NH 活性的骨架)。
- 第四,二维实验(COSY、HSQC、HMBC)和立体化学问题在设计上被排除在外,因为仅凭一维核磁共振无法确定构型。因此,复杂的天然产物化合物未纳入评估。
- 最后,我们的溶剂覆盖范围仅限于 DMSO-d₆、CDCl₃ 和 D₂O,因此甲醇-d₄、苯-d₆ 和丙酮-d₆ 未进行评估。
理想情况下,我们希望看到这些数据在横跨 20 到 30 种骨架类别的数百种化合物上的表现,每个类别至少包含 15 种化合物,以便将类别内的差异与不同工具间的差异区分开来。我们还会评估氯哒嗪之外的含 NH 活性杂芳烃,评估未测试的溶剂,并开展利用二维实验的两种任务版本。
未来展望
在我们持续提升 Claude 在化学领域性能的过程中,我们特别聚焦于几个最拖慢化学家工作效率的瓶颈问题。
- 化学结构的读取与呈现——将来自图表、专利、幻灯片或手绘草图中的结构式转换为机器可读的形式,并在结构式表示与化学文献中使用的系统命名之间进行转换。
- 反应与合成推理——提出、评估和评判合成路线,预测反应结果,并思考选择性、反应条件以及可能的副产物。
- 机理——以化学家实际使用的语言(包括电子箭头、中间体和过渡态论证)来解释和检验反应机理。
- 化学文献理解——阅读已发表作品中的化学内容,其中同一分子可能以结构式、名称、缩写或代码形式出现,并从方法部分、补充信息和专利中提取出关键化学信息。
这些方面并非都处于相同的成熟度曲线上。光谱分析已经足够成熟可以进行基准测试,而其他方面,如逆合成规划,仍处于范围界定阶段。随着我们对这些瓶颈问题有了更深入的了解,我们将分享当前模型在哪些方面表现出色,以及在哪些方面仍有不足。我们的最终目标是确保一线化学家知道 Claude 能在哪些方面为他们节省时间,以及在哪些方面他们仍需依赖自身的专业知识。
与我们合作
我们正在扩展“AI 助力科学”项目,以更明确地支持化学研究。如果您是正在研究某个问题、且 Claude 有可能提供帮助的研究人员,尤其是涉及我们描述的那类多模态推理的问题,我们期待您通过 scienceblog@anthropic.com 或“AI 助力科学”申请页面与我们联系。
脚注
- 一起与一种孕期止吐药物相关、在全球导致超过一万名儿童出现严重出生缺陷的事件。
- 我们从中提取化合物的四篇预印本:https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002274/v1,https://chemrxiv.org/doi/full/10.26434/chemrxiv-2025-59lfh,https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002423/v1,https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002316/v1。
加拿大如何使用 Claude:来自 Anthropic 经济指数的发现
Claude 在不同模型和语言中的价值观
We’re working with world-class synthetic, computational, and analytical chemists to make Claude better at chemistry. In this post, we share our first work as part of this effort, in which Anthropic chemist, David Kamber, examines how Claude performs on a chemist’s most common analytical input, an NMR spectrum.
When working with molecules, chemists move between hand-drawn structures on a whiteboard, instrument readouts, database query strings, and the technical notations of patents and publications. Each of these representations encodes the same underlying chemistry, but each demands a different kind of fluency. A sketch of caffeine, for example, allows a chemist to spot its resemblance to adenosine, the body’s drowsiness signal, and predict that it keeps us alert by blocking the receptor. However, that same sketch cannot help a chemist tell it apart from other near-identical looking molecules.
Understanding what molecule a chemist is working with is critical. Chemistry undergirds everything from the foods and medicine we ingest to our lotions, paints, and plastics. Reroute a handful of bonds among the same atoms, and glucose becomes fructose, molecules sharing a formula but processed through entirely different metabolic pathways. Flip a molecule into its mirror image, and a sedative becomes a teratogen, as happened in the thalidomide disaster.1 Chemists’ everyday work depends on reading these signals correctly across whichever representation befits a given task.
Translating between these representations (chasing down a structure from a figure, reconciling an instrument readout against a proposed product, querying a database in the right notation) is time consuming and impossible to keep up with at scale—CAS, the largest chemistry registry, catalogs over 290 million disclosed substances and grows by roughly 15,000 new ones every day.
AI is well-positioned to take on this research burden, yet it still remains largely aspirational in the context of chemistry. Machine-learning tools have been positioned for years as transformative for retrosynthesis—the process of working backward from a target molecule to simpler precursors to plan how to build it—reaction prediction, and property estimation, but the data those tools need have been hard to come by—sparse on null-results, inconsistent in format, and locked behind paywalls at subscription journals (and in unstructured supporting information). Retrosynthesis is a case in point—capable AI tools have existed for years, but adoption is uneven, and the average academic or small-lab chemist still doesn't use them.
Even so, advancements in AI are finally reaching chemistry. Today’s frontier models are multimodal, and capable of explicit reasoning. They can read a chemical structure directly from a journal figure or hand sketch rather than depending on a pre-curated molecular database. And they can read the experimental detail of a methods section or supporting information in the form it is actually published. They can also show their reasoning step by step, which means a chemist can audit the outputs. None of this eliminates the data problem the field has been describing for years, but it changes which problems are tractable despite it.
Ultimately, our claim is a modest one: Claude is starting to meaningfully assist chemists with the daily translation, recall, and integration work that complements their judgment, and we plan to keep extending its helpfulness. Today we are publishing the first white paper in the effort to accelerate this work. It tackles a chemist's most common analytical input: an NMR spectrum.
Claude vs. ChemDraw on NMR prediction and structure elucidation
Full version can be found here
Nearly every small molecule—drug, pesticide, dye, fragrance, polymer, DNA or protein subunit, and functional inorganic or solid-state material—exists because a chemist determined its structure. Given that these molecules cannot be seen with microscopes, chemists must rely on spectral analysis, probing a molecule with light, radio waves, or magnetic fields. The way a given molecule absorbs, emits, or deflects this energy gives chemists a pattern, or spectrum, with which they can elucidate its structure.
NMR spectroscopy—one of the canonical techniques chemists rely on for this—is one of the most time-consuming steps in synthetic chemistry; for every compound, a chemist has to match each peak in the spectrum to an atom in the proposed structure by hand. For this white paper, we tested how Claude fared against the dedicated NMR software chemists rely on today. We measured three Claude models (Opus 4.7, Opus 4.6, Sonnet 4.6) against ChemDraw and MestReNova on 20 compounds drawn from synthetic chemistry preprints published after the models’ training cutoff so as to avoid selection bias. Both ChemDraw and MestReNova do forward prediction, using a drawn structure to simulate what NMR spectrum will be produced. In addition to forward prediction, we also wanted to see whether Claude could go the other direction—starting from an experimental spectrum and proposing the structure behind it. This is the harder task, and the one existing software currently leaves to the chemist.
To set up our assessment, we pulled 20 compounds from ChemRxiv preprints2 posted after the models’ training cutoff, taking the first fully characterized novel molecules from each paper. The 20 span four structural families, five compounds each, with each family selected because it involves a different category of NMR challenge. Each tool was given the structure encoded as a SMILES string—the line-of-text notation chemists use to input a molecule to software—and was asked to predict where every hydrogen and carbon peak would fall along a 1D NMR spectrum (a horizontal axis measuring chemical shifts in ppm, parts per million). Given that NMR samples are dissolved in a liquid, and that the choice of solvent (chloroform, DMSO, etc.) moves the peak positions slightly, each tool was told to predict the spectrum in whatever solvent the chemists had used in the published paper.

Because a language model’s output varies between runs, each Claude model was queried three times per compound and averaged; ChemDraw and MestReNova return the same answer every time and were run once. We then paired each predicted peak with its experimental counterpart and measured the gap in ppm. These landed within the window a chemist would call correct—±0.20 ppm for hydrogen or ±1.0 ppm for carbon.

On hydrogen, Opus 4.7 was most accurate, with an average error of ±0.079 ppm—well under half the tolerance window—and the highest share of peaks landing inside it. On carbon, Opus 4.7 and MestReNova were effectively tied, at ±1.37 and ±1.48 ppm; the remaining tools kept the same rank order on both elements. Opus 4.6 was predictably middling, and Sonnet 4.6 was the weakest. The gap between them was most evident on a single notoriously difficult hydrogen—an NH proton in the chloropyridazine family whose true position falls in a narrow band between 6.8 and 7.9 ppm. Opus 4.7 placed it slightly low but consistently so; Opus 4.6 scattered its guesses across several ppm; Sonnet 4.6 put it in the 10–13 range, well outside where it actually appears.

While Opus 4.7 performed fairly comparably to ChemDraw and MestReNova, the gap was wider on predicting the shape taken by a hydrogen’s NMR peak and how far apart the peaks sit, features which also contain structural information a chemist reads alongside position. Opus 4.7 matched the experimentally reported splitting pattern more often than any other tool, and all three Claude models predicted the sub-peak spacing to within half a hertz roughly 80% of the time—against 26 to 35% for ChemDraw and MestReNova. Opus 4.7 was also the most consistent across its three repeat runs: its average error varied less from run to run than the margin separating it from the next-best tool.
From there, we evaluated inverse prediction (structure elucidation): could we determine the structure of a molecule from its spectrum? We gave Opus 4.7 15 elucidation problems and asked it, three times each, to propose up to three ranked candidate structures. Each supplied the compound’s exact molecular formula (from high-resolution mass spectrometry) and its hydrogen and carbon NMR spectra. The fifteen were split by difficulty. The eight simpler targets—single-ring or two-fragment molecules—were posed with only the formula and spectra. The seven denser targets—fused rings, spirocycles, and similar—were accompanied by one additional hint: the structure of the starting material that had gone into the reaction.

Opus 4.7 recovered all eight simpler structures on every attempt from spectra and formula alone. On the seven harder targets, given the starting-material hint, it returned the correct structure on all three runs for four of them and on two of three runs for those that remained.
Ultimately, we found that for routine data prediction Opus 4.7—a general-purpose model without chemistry-specific fine-tuning—is now as good as or better than ChemDraw and MestReNova on average. Additionally, Claude can also work the problem in reverse, proposing a structure from NMR data alone. Dedicated structure-elucidation software has existed for decades, but it typically requires 2D NMR (a spectrum with two axes, and the output is a contour map rather than a row of peaks), specialized training, and licensed tools. Claude does it from the same high-resolution mass spectrum and 1D peak list a chemist would paste into a chat, with no setup.
Limitations
This assessment shows us that a general-purpose model can be competitive with NMR software and even make 1D inverse elucidation tractable. But there are a handful of noteworthy limitations.
- First, the evaluation was small—20 compounds across four scaffolds for the forward task, 15 for the inverse task—and each scaffold contributes a single class of failure modes. The model performance should thus be read as indicative rather than precise.
- Second, on the densest inverse targets, without the starting material as an additional input, the model could loop through its reasoning without committing to a final structure; this is why the seven harder problems were posed with the starting-material structure rather than spectra alone.
- Third, some chemical scaffolds were left untested. For example, slow-exchange NH heteroaromatics (aromatic rings whose N–H exchanges with solvent slowly enough to leave a sharp NMR peak) are sampled only through chloropyridazines, leaving out related systems (hydroxypyridines, aminothiazoles, and other DMSO-d₆ NH-active scaffolds).
- Fourth, 2D experiments (COSY, HSQC, HMBC) and stereochemistry are out of scope by design, since 1D NMR alone cannot fix configuration. As a result, complex natural product compounds were not evaluated.
- And finally, our solvent coverage was limited to DMSO-d₆, CDCl₃, and D₂O, so methanol-d₄, benzene-d₆, and acetone-d₆ are not assessed.
Ideally, we would see how these numbers hold up across several hundred compounds spanning 20–30 scaffold classes, with at least 15 compounds per class so that within-class variance can be separated from between-tool differences. We would also evaluate NH-active heteroaromatics beyond chloropyridazines, assess the untested solvents, and conduct versions of both tasks that draw on 2D experiments.
Looking ahead
As we continue to improve Claude’s performance in chemistry, we are focusing specifically on a handful of bottlenecks that slow chemists down the most.
- Reading and rendering chemical structures—converting a drawing from a figure, patent, slide, or sketch into a machine-readable form, and going between structural representations and the systematic names used in chemistry literature.
- Reaction and synthetic reasoning—proposing, evaluating, and critiquing synthetic routes, anticipating outcomes, and thinking through selectivity, conditions, and likely byproducts.
- Mechanism—explaining and testing reaction mechanisms in the language a chemist actually uses, with electron arrows, intermediates, and transition-state arguments.
- Chemical literature understanding—reading chemistry as it appears in published work, where the same molecule may be drawn, named, abbreviated, or referenced by a code, and pulling out the chemistry that matters from method sections, supporting information, and patents.
These are not all on the same maturity curve. Where spectral analysis is far enough along to benchmark, others, like retrosynthesis planning, are still being scoped. As we get a better understanding of these bottlenecks, we will share where current models excel, and where they still fall short. Our ultimate goal is to ensure that working chemists know where Claude can save them time and where they still need to rely on their own expertise.
Working with us
We are expanding the AI for Science program to more explicitly support chemistry research. If you are a researcher working on a problem where Claude could plausibly help, especially one that involves the kinds of multimodal reasoning we have described, we would like to hear from you at scienceblog@anthropic.com, or through the AI for Science application.
Footnotes
- An incident in which a morning sickness medication was linked to severe birth defects in over 10,000 children worldwide.
- The four preprints from which we pulled the compounds: https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002274/v1, https://chemrxiv.org/doi/full/10.26434/chemrxiv-2025-59lfh, https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002423/v1, https://chemrxiv.org/doi/full/10.26434/chemrxiv.15002316/v1.