# AI 药物研发实际进展如何：临床影响证据仍有限

- 来源：Hacker News 热门（buzzing.cc 中文翻译）
- 作者：AnodicElegy
- 发布时间：2026-08-16 10:46
- AIHOT 分数：48
- AIHOT 链接：https://aihot.virxact.com/items/cmsv7lw2w0ae5rosaxdkvargu
- 原文链接：https://www.science.org/content/blog-post/so-how-ai-drug-discovery-doing-really

## AI 摘要

多位领域专家发文指出，尽管各类 AI 方法已被开发、应用和基准测试，其临床相关影响的证据迄今仍“令人失望地有限”，目前处于“缺乏证据”阶段。文章强调，AI 在药物研发中的重点应从“做能做的事”转向“做应做的事”，并指出 II 期临床成功率是衡量改进的关键指标，但现有数据因混杂因素和条件性而难以用于真正的机器学习。

## 正文

So How Is AI Drug Discovery Doing, Really?

10 Aug 2026

3:34 PM ET

By Derek Lowe

2 min read

Comments

The deluge of articles, presentations, and press releases on the applications of AI techniques to drug discovery does not make it easy to tell just what those applications are and how useful they might be (at least so far). This is an excellent new article from a number of people who know what they’re talking about in the field and who (importantly) are not trying to sell you anything, either.

Its authors point out (correctly, I’d say) that the evidence for real clinical impact so far is very thin:

“Although a wide variety of AI methods have been developed, applied and benchmarked, evidence of their clinically relevant impact is, so far, disappointingly limited. . .It can be concluded that we are still at a stage with an ‘absence of evidence’ (and not necessarily an ‘evidence of absence’) when it comes to the translation of AI into clinically relevant impact.”

One of the effects you’d really like to see is better decision making with the help of AI - better selection of disease areas, of targets, of lead molecules and clinical candidates, of trial designs and more. I go around telling people that these are some of the areas that current AI techniques are least equipped to help us out with, but I always need to emphasize that world “current”. There’s no fundamental reason why these things can’t be improved, but such improvements are going to be slower and more expensive than claimed by many of the people who send out all those press releases.

You can see why it’s hard to get good data on these effects. Even without AI, it’s been hard! There’s a natural tendency to ascribe successful projects to keen insights, bold leadership, and relentless hard work, while unsuccessful ones, umm, is there some reason why you’d want to talk about those? I mean, really? Because if you buy into the story completely, that means that unsuccessful projects were doomed by what, fuzzy thinking and laziness? I’m not saying that those won’t doom your project - they probably will - but they’re not the only reasons that drug discovery and development efforts fail. There are plenty of examples where people did the right things the right way and for the right reasons and still found themselves sliding into the mudpit. So you get the “When we win we’re England, when we lose we’re MCC” effect (deep British sports history cut there). Any project that had any AI aspect to it likely succeeded because of that, don’t you know, and the other projects that failed even though they had the Digital Laying On of Hands, well, moving right along. . .

The paper rightly identifies Phase II success rates as where you really would want to see the most impact from any purported improvements in drug development. All that stuff in early-stage discovery work is a roundoff error in time and money compared to the clinical trials, and Phase II is where you first run into the really large cross-section for failure in the clinic. This has of course not escaped anyone’s notice in this business, and pre-AI there have been all sorts of Capitalized Revolutionary Acronymic Protocols announced in various companies to try to get those success rates up. None of this has had much impact in the end. The AI era has just been more of the same, as far as I can tell so far.

The paper exhorts practitioners to get down to it: “. . .the focus of AI in drug discovery must shift from doing what can be done - such as modelling data that is readily available, but that is unlikely to move the needle - to doing what should be done, even if this requires, for example, substantial data generation. . .” It’s a worthy goal, but I think that many involved in this work might be thinking, even unconsciously, “You first”. Making this commitment is an expensive proposition with murky chances of success (albeit for very high rewards), and it also implicitly means telling the investors (and the press) that things are going to take longer than they’re used to hearing about.

There’s a very useful discussion of the difference between a ligand and a drug, which all to often has translated to the difference between the preclinical organization’s goals and the clinical ones that come afterwards. As always, you should work in as close to the real system as you can: if you say you have a good compound in a cloned isolated protein, I want to know how it is in cells. If it works in cells, tell me about the rodents. If it works there, I want to know about two-week dog tox, and if it gets through that, then why aren’t you putting it into human patients now? Finding AI systems that will move those steps along and (far more importantly) make them translate into clinical success is not easy.

The authors also go into another major problem: sure, we have piles and piles of data about compounds in all sorts of assays (including all those stages just mentioned). But machine-learning your way to happiness with those numbers is very likely impossible, due to the huge numbers of variables in the assay conditions and indeed, in biological systems themselves. We don’t know how to clean this stuff up and categorize it for real ML/AI, and honestly, we don’t even know if it can be. The same considerations apply to data that you might hope would be less shaggy (structure-activity relationships in the chemistry, for example). This is a good point on top of all of these:

Another key problem is that the nature of the data at hand in drug discovery — its confounding factors, conditionality and epistemic opacity — may not be understood by everyone aiming to model such data, and hence the limitations of labelling data and using labelled data are not understood either. As a consequence of ‘believing the labels’, there is a considerable knock-on effect on benchmarks that are being used in the field. Benchmarks have been crucial to provide some estimate of progress in AI, given that otherwise no quantitative comparison of novel approaches to the existing ‘state-of-the-art’ (SOTA) would be possible. In some areas, such as image recognition and speech recognition, benchmarks by and large translated to better performance in real-world use cases. However, in the area of data used in drug discovery, this has to be viewed with caution

The paper goes on to make recommendations for AI companies and investigators, and these are well worth reading. The common theme is that people need to think more about why they’re doing certain techniques or using certain technologies, rather than just using them because they’re newly available. How many of these can (realistically) improve the chances of clinical success? Do you have any good ways to estimate that? What would you need to get such estimates? Are you guarding at every stage against the tendency to take shortcuts or to believe your own hype? Are you making sure that you’re not setting up benchmarks and targets that are better designed to give you (and others!) the illusion of progress rather than delivering the real thing? Do you have enough time and money to get to the real thing at all?

Real application of the principles laid out in this paper would do the field a lot of good, probably save a number of investors some very unpleasant experiences, and would absolutely turn down the noise of the hype machines. Not everyone is aligned on those goals, however. . .

About the author

Derek Lowe

Comments

Listen as Science staff delve into scientific discoveries and policy with researchers and writers from around the globe.
