Gary Marcus:The Road to AI We Can Trust(RSS)
54AI 编辑部评分,满分 100

大语言模型(LLMs)是否改善了患者治疗结果?

2026-05-04 04:00· 93天前· Gary Marcus
AI 导读

一项新综述研究指出,尽管大语言模型(如GPT、Claude、LLaMA)在医疗领域的应用日益广泛,但目前尚无明确证据表明其直接改善了患者治疗结果。该综述分析了多项临床研究,发现这些模型在诊断支持、文书处理等方面展现出潜力,但在提升治愈率、降低死亡率或改善患者生活质量等关键临床指标上,尚未展现出统计学上的显著积极影响。研究强调,需要更多高质量的随机对照试验来评估LLMs对患者结局的实际影响。

A new review suggests otherwise

Artwork by Dave Coverly at Speedbump.com. Reprinted by perrmission.

Cardiologist/author Eric Topol tends to be more bullish than I am about AI in medicine, but in his latest post he concludes that at least thus far there is “very little evidence for LLMs benefiting patients or doctors for health outcomes” (aside from helping with administrative work and the like).

As Topol notes, his review was partly inspired by thisrecent editorial in Nature Medicine, which points in the same direction:

You can read Topol’s detailed review here.

All of this fits with my recent Please don’t trust your chatbot for medical advice, which focused more on unsupervised use of LLMs by patients, rather than overall clinical outcomes, but in many ways converges to a similar place.

AI will surely someday be a major boon for medicine, but current tools such as domain-general chatbots may not be up to the job.

Discussion about this post

The Topol finding is interesting precisely because hes more bullish than you and still arrived at the same conclusion. When the optimist and the sceptic converge on "not yet," thats a stronger signal than either position alone.

The administrative work exception is the tell though. LLMs are genuinely good at summarising notes, drafting letters, handling paperwork — tasks where being wrong has low consequences and being fast has high value. The moment you move into clinical decision-making where being wrong means someone gets hurt, the accuracy threshold jumps from "good enough" to "better than the doctor" and thats a completely different bar.

The gap isnt really about whether LLMs can pass medical exams. They can. Its about whether they fail gracefully. A doctor who isnt sure orders more tests. An LLM that isnt sure sounds exactly as confident as one that is. And in medicine, the difference between uncertainty expressed and uncertainty hidden is sometimes the difference between a patient who lives and one who doesnt.

The amount of money wasted by AI in the past two years could have been better utilized to improve healthcare, leading to better outcomes.

Ready for more?

来源:Gary Marcus:The Road to AI We Can Trust(RSS) · garymarcus.substack.com

大语言模型(LLMs)是否改善了患者治疗结果?

Gary Marcus:The Road to AI We Can Trust(RSS)·2026-05-04 04:00·93天前·Gary Marcus
AI 导读

一项新综述研究指出,尽管大语言模型(如GPT、Claude、LLaMA)在医疗领域的应用日益广泛,但目前尚无明确证据表明其直接改善了患者治疗结果。该综述分析了多项临床研究,发现这些模型在诊断支持、文书处理等方面展现出潜力,但在提升治愈率、降低死亡率或改善患者生活质量等关键临床指标上,尚未展现出统计学上的显著积极影响。研究强调,需要更多高质量的随机对照试验来评估LLMs对患者结局的实际影响。

原文 · 保持原样,未翻译

A new review suggests otherwise

Artwork by Dave Coverly at Speedbump.com. Reprinted by perrmission.

Cardiologist/author Eric Topol tends to be more bullish than I am about AI in medicine, but in his latest post he concludes that at least thus far there is “very little evidence for LLMs benefiting patients or doctors for health outcomes” (aside from helping with administrative work and the like).

As Topol notes, his review was partly inspired by thisrecent editorial in Nature Medicine, which points in the same direction:

You can read Topol’s detailed review here.

All of this fits with my recent Please don’t trust your chatbot for medical advice, which focused more on unsupervised use of LLMs by patients, rather than overall clinical outcomes, but in many ways converges to a similar place.

AI will surely someday be a major boon for medicine, but current tools such as domain-general chatbots may not be up to the job.

Discussion about this post

The Topol finding is interesting precisely because hes more bullish than you and still arrived at the same conclusion. When the optimist and the sceptic converge on "not yet," thats a stronger signal than either position alone.

The administrative work exception is the tell though. LLMs are genuinely good at summarising notes, drafting letters, handling paperwork — tasks where being wrong has low consequences and being fast has high value. The moment you move into clinical decision-making where being wrong means someone gets hurt, the accuracy threshold jumps from "good enough" to "better than the doctor" and thats a completely different bar.

The gap isnt really about whether LLMs can pass medical exams. They can. Its about whether they fail gracefully. A doctor who isnt sure orders more tests. An LLM that isnt sure sounds exactly as confident as one that is. And in medicine, the difference between uncertainty expressed and uncertainty hidden is sometimes the difference between a patient who lives and one who doesnt.

The amount of money wasted by AI in the past two years could have been better utilized to improve healthcare, leading to better outcomes.

Ready for more?

来源:Gary Marcus:The Road to AI We Can Trust(RSS)· garymarcus.substack.com