医信观察 · MED IT
中文译文数据治理与互联互通

更好的模型无法解决制药业的 AI 问题——更好的术语可以

Better Models Won’t Fix Pharma’s AI Problem — Better Terminology Will

MedCity News··约 7 分钟阅读
译文3,344 字

今年早些时候,在一次生命科学会议上,一位制药分析负责人把我拉到一旁,问了一个听起来很简单的问题:“我们正在研究的疾病有 ICD-10 编码吗?”没有。正如许多研究人员所知,现实情况是,数千种医学状况和疾病状态都存在这一问题。当一种状况本身无法在数据中被清晰捕获和索引时,建立在这些数据之上的每一项 AI 分析或决策都会开始摇摆不定。

AI 驱动的真实世界数据分析的瓶颈,不在于计算能力、模型架构或训练数据量,而在于 AI 试图据此进行推理的语义层——精确临床含义所在之处。

由以下机构呈现赞助文章弥合职场心理健康领域的质量与可负担性鸿沟在一次采访中,Kyan Health 联合创始人兼首席商务官 Konstantin Struck 讨论了 Kyan 如何让中型市场和大型企业雇主以可负担的价格获得优质的员工心理健康护理。

作者:Stephanie Baum当今数据仍无法提出的问题大多数用于真实世界数据分析的 AI 系统都建立在 ICD-10、SNOMED CT 和 LOINC 等标准代码集之上。这些标准在计费、文档记录和互操作性方面实现了一致性,但它们并不是为现代制药团队试图回答的研究问题而设计的。这一挑战可以通过三个常见的失效点来说明。

罕见病队列:

在许多情况下,罕见病患者实际上在数据集中是不可见的,因为没有精确的代码可以识别他们。例如,Orphanet Journal of Rare Diseases 于 2024 年发表的一项研究发现,454 种罕见病中只有 34% 能够在 ICD-10-GM 中进行特异性编码。

严重程度梯度和疾病活动度:临床医生和患者都非常重视疾病早期与晚期,或轻度、中度与重度表现之间的差异。这些具有临床意义的差异在用于研究和患者队列划分的代码集中往往会被压缩为单一的标准化类别,因此更难分离出研究人员所关注的患者亚组。

赞助文章

AI 部署将在哪些方面最有利于付款方?

我们正在了解健康保险公司如何使用 AI、如何定义成功,以及如何管理网络安全风险。请完成我们的简短匿名调查,告诉我们您的看法。

作者:MedCity News

表型分析和亚型划分:精准医疗越来越依赖分子和临床层面的特异性,而计费时代的术语系统从未被设计用来承载这种特异性。这种特异性往往隐藏在电子健康记录(EHR)的自由文本部分中。JMIR Medical Informatics 发表的一项研究显示,从患者记录中提取的概念只有 13%(从就诊记录中提取的概念只有 7%)在结构化代码与自由文本记录之间存在任何重叠,这表明绝大多数临床信息只存在于其中一种形式,而非两者兼有。

我花费大量时间与正在努力应对这些问题的制药业领导者交流——他们试图利用最好的 AI 工具来增强工作,却受到远未达到适用目的的数据的掣肘。令人震惊之处在于这一挑战的规模和普遍性。关键的临床细微差别往往在抵达 AI 模型之前就已经消失。结果是可以预料的:在演示中看起来很复杂的模型,在临床医生和研究人员真正需要得到答案的问题上却会失效。

队列由可计费的内容定义

实际情况相对简单。生命科学分析中使用的大多数 AI 模型都基于已按照可计费代码集标准化的数据进行训练,因为这正是行业能够获得的数据。源自理赔的术语主导着许多真实世界数据管线,即使这些数据通过 EHR 系统中的临床数据得到增强,临床保真度也往往会丢失。

这种动态造成了一种微妙但影响深远的扭曲:队列由可计费的内容定义,而不是由临床事实定义。

考虑一个常见的研究设计问题。研究人员希望识别患有轻度、中度或重度疾病的患者。标准化代码集可能只写着“L40:银屑病”。但临床医生可能记录的是“银屑病,中度严重程度,伴有类风湿性关节炎共病”。

正是这种特异性决定了试验入选资格、治疗反应和后续结局。这种细微差别存在于医疗服务点,却在进入理赔层的瞬间被压平。

Frontiers in Digital Health 发表的一项疫苗接种研究展示了这种“压平”的影响。研究人员利用自然语言处理(NLP)分析患者记录中的非结构化数据,发现与仅使用结构化数据相比,NLP 使疫苗接种记录的识别率提高了 16.8%,突出了仅依赖结构化 EHR 数据的局限性。

同样的模式也出现在肿瘤学和罕见病研究中。分子亚型、疾病进展标志物以及具有临床意义的修饰因素,往往会消失在更宽泛的管理类别中,而这些类别从未被设计用于支持精准研究。

用于研究的 AI 推理模型随后继承了这种扁平化。研究人员不明白为什么队列选择会纳入实际上不符合研究入选标准的患者,而根本问题在于,源术语从未保留临床医生最初作出的那些区分。

这是精确性问题,而不是规模问题

行业对 AI 局限性的回应,很大程度上是追求更大的模型和更大的训练集。然而,每一项具有临床意义的 AI 决策都依赖于更基础的一点:模型底层的数据是否能够准确表达医生在记录患者就诊时的意图。

这种保真度会影响临床试验中的患者识别、文献证据综合、真实世界队列划分,以及由真实世界证据支持的监管申报。如果临床含义在上游被稀释,再高明的下游建模也无法完全恢复它。

这就是为什么围绕可信医疗 AI 的讨论越来越少关注模型本身,而更多关注其底层的临床智能。

AI 能够可靠据此进行推理的术语层具有几个定义性特征:

首先,它必须由医生、术语学家和主题专家进行整理和临床验证,而不是被动汇编,或仅凭蛮力从计费数据中进行机器学习。

其次,它必须具备溯源能力。研究人员应能够将概念追溯至可信的临床来源,并理解某位患者为何被纳入某个队列。

最后,它必须作为一个相互连接、不断演进的知识图谱运行,而不是静态的查找表,使研究人员能够在 EHR 数据、登记库、理赔数据和文献之间保持一致地进行推理,同时不丢失保真度。

这正是许多医疗 AI 部署仍然缺失的基础设施层。这从根本上说不是模型问题,而是精确性问题。在医疗行业最无力承担这种代价的时刻假装并非如此,可能会侵蚀人们对医疗 AI 的信任。

对于生命科学研究负责人而言,其影响是立竿见影的。审查 AI 系统下方的数据层。询问模型据以进行推理的术语和临床内容是什么、端到端有多少特异性得以保留,以及这些特异性在哪里丢失。

要求提供溯源能力,确保每项队列定义和证据综合都能够追溯到临床源材料。在投资于另一轮模型升级之前,评估现有系统底层的术语是否保留了你所需要的临床含义。

在术语不足的情况下,更先进模型的边际价值相对较小;在一个能力合格的模型下方采用更好的术语,其边际价值则极其巨大。

对具有临床意义的数据的追求

我一直在回想会议上的那次谈话。那位研究人员的问题并不复杂。他们只是想知道,在真实世界数据和临床实践中,是否存在用于识别其所研究患者的语言。借助精准术语,这个问题就可以得到回答。没有精准术语,即使使用最好的 AI 工具加以增强,整个研究类别仍会对分析隐藏不见。

医疗领域的 AI 不会因为模型不够“聪明”而失败。它仍将受到限制,因为我们使用无法完整表达临床医生意图的数据训练了这些系统。

未来十年在药物发现、商业化和患者可及性方面处于领先地位的组织,将是在迭代其 AI 工具下一版本之前,确保数据底层临床含义忠实性的组织。

图片:claudenakagawa,Getty Images Joseph Zabinski Joseph Zabinski,PhD、MEM,于 2025 年加入 IMO Health,担任产品管理高级副总裁。他负责生命科学产品组合和上市战略。Zabinski 博士是利用医疗数据开展 AI 驱动的个性化和结局预测领域的著名作者及思想领袖。此前,他负责 OM1 的 AI 与个性化医疗业务部门;OM1 是一家真实世界证据和技术提供商。Zabinski 博士的职业生涯始于 McKinsey 的顾问工作,为制药客户提供 AI 战略和实施方面的建议。他在 UNC Chapel Hill 的 Gillings School of Global Public Health 获得博士学位,研究重点是应用于大型医疗数据集的贝叶斯图模型。他还获得了 Dartmouth 的工程管理 MEM 学位、Boston College 的物理学和德语研究 BS 学位,以及赴奥地利的 Fulbright Fellowship。

本文通过 MedCity Influencers 项目发布。任何人都可以通过 MedCity Influencers,在 MedCity News 上发表自己对医疗领域商业和创新的观点。

点击此处了解具体方式。

原文8,205 字符

At a life sciences conference earlier this year, a pharma analytics lead pulled me aside with what sounded like a simple question: “Is there an ICD-10 code for the disease we’re studying?” There wasn’t. As many researchers know, the reality is that this is true for thousands of medical conditions and disease states. When a condition itself cannot be captured and indexed cleanly in the data, every AI-enabled analysis or decision built on top of that data starts to wobble.

The bottleneck in AI-enabled real-world data analysis is not computation, model architecture, or training data volume. It is the semantic layer over which the AI is trying to reason – the place where precise clinical meaning lives.

The questions today’s data still can’t ask

Most AI systems in real-world data analysis are built on standard code sets like ICD-10, SNOMED CT, and LOINC. Those standards create consistency across billing, documentation, and interoperability, but they were not designed for the kinds of research questions modern pharma teams are trying to answer. This challenge can be illustrated through three frequent failure points.

Rare disease cohorts:

In many cases, patients with rare diseases are effectively invisible in datasets because there is no precise code to identify them. For example, a 2024 study in Orphanet Journal of Rare Diseases found that just 34% of 454 rare diseases could be specifically coded in ICD-10-GM.

Severity gradients and disease activity

: Clinicians and patients care deeply about distinctions between early- and late-stage disease, or mild versus moderate versus severe presentation. These clinically meaningful differences often collapse into a single standardized category in codesets used for research and patient cohorting, making it more difficult to isolate patient subgroups of interest.

Sponsored Post

Where Will AI Deployment Benefit Payers the Most?

We are taking a look at how health insurers are using AI, defining success, and managing cybersecurity risks. Give us your opinions by completing our brief, anonymous survey.

By MedCity News

Phenotyping and subtyping: Precision medicine increasingly depends on molecular and clinical specificity that billing-era terminology systems were never designed to carry. That specificity tends to hide in the free-text section of electronic health records (EHRs). A study in JMIR Medical Informatics showed that only 13% of extracted concepts from patient records (and 7% from visits) showed any overlap between structured codes and free-text notes, indicating that the vast majority of clinical information exists in one form or the other, not both.

I spend much of my time talking with pharma leaders wrestling with these issues – trying to use the best AI tools to enhance their work, but hobbled by data far from fit-for-purpose. What makes the challenge striking is its scale and ubiquity. Critical clinical nuance routinely disappears before it ever reaches an AI model. The result is predictable. Models that look sophisticated in demos fail at the questions clinicians and researchers actually need answered.

The cohort gets defined by what was billable

What is happening on the ground is relatively straightforward. Most AI models used for life sciences analytics are trained on data normalized to billable code sets because that is the data the industry has available. Claims-derived terminology dominates many real-world data pipelines, and even when enhanced with clinical data from EHR systems, clinical fidelity is frequently lost.

This dynamic creates a subtle but consequential distortion: The cohort gets defined by what was billable, not what was clinically true.

Consider a common study design problem. A researcher wants to identify patients with mild, moderate, or severe disease. The standardized code set may simply say “L40: psoriasis”. But the clinician may have documented “psoriasis, moderate severity, with comorbid rheumatoid arthritis.” That specificity is exactly what determines trial eligibility, treatment response, and downstream outcomes. The nuance exists at the point of care, only to be flattened the moment it enters the claims layer.

A study of vaccine administration published in Frontiers in Digital Health illustrates the effects of this “flattening.” Using natural language processing (NLP) to analyze unstructured data in patient records, researchers found that NLP led to a 16.8% increase in the identification of vaccine administrations compared with using structured data alone, highlighting the limitations of relying on structured EHR data alone.

The same pattern appears in oncology and rare disease research. Molecular subtypes, progression markers, and clinically meaningful modifiers often disappear into broader administrative categories that were never intended to support precision research.

AI reasoning models for research then inherit this flatness. Researchers wonder why cohort selection yields patients who are not truly study-eligible, when the underlying issue is that the source terminology never preserved the distinctions that clinicians made in the first place.

This is a precision problem, not a scale problem

The industry response to AI limitations has largely been to chase larger models and larger training sets. However, every clinically meaningful AI decision depends on something more fundamental: whether the data beneath the model can accurately express what a physician meant when documenting a patient encounter.

That fidelity affects patient identification for clinical trials, evidence synthesis from literature, real-world cohorting, and regulatory submissions backed by real-world evidence. If the clinical meaning is diluted upstream, no amount of downstream modeling sophistication can fully recover it.

This is why the conversation around trustworthy healthcare AI is increasingly less about the model itself and more about the clinical intelligence beneath it.

A terminology layer that AI can reliably reason over has several defining characteristics:

First, it must be curated and clinically validated by physicians, terminologists, and subject matter experts, not assembled passively or machine-learned by brute force from billing data alone.

Next, it must be provenanced. Researchers should be able to trace concepts back to trusted clinical sources and understand why a patient was included in a cohort.

Finally, it must function as a connected, evolving knowledge graph rather than a static lookup table, allowing researchers to reason consistently across EHR data, registries, claims, and literature without losing fidelity.

That is the infrastructure layer that many healthcare AI deployments are still missing. This is not fundamentally a model problem. It is a precision problem. Pretending otherwise risks eroding trust in healthcare AI at exactly the moment the industry can least afford it.

For life science research leaders, the implications are immediate. Audit the data layer underneath your AI systems. Ask what terminology and clinical content the model is reasoning over, how much specificity survives end-to-end, and where that specificity gets lost.

Demand provenance so every cohort definition and evidence synthesis can be traced back to clinical source material. And before investing in another model upgrade, evaluate whether the terminology underlying your existing systems preserves the clinical meaning you need.

The marginal value of a more advanced model, given inadequate terminology, is relatively small. The marginal value of better terminology underneath a competent model is enormous.

The quest for clinically meaningful data

I keep thinking back to that conversation at the conference. The researcher’s question was not complicated. They simply wanted to know whether the language existed in real-world data and clinical practice to identify the patients they were studying. With precision terminology, that question becomes answerable. Without it, entire categories of research remain obscured to analysis, even when enhanced with the best AI tools.

AI in healthcare will not fail because the models are insufficiently “smart.” It will remain limited because we trained these systems on data that could not fully express what the clinician meant.

The organizations that lead the next decade of drug discovery, commercialization, and patient access will be the ones that ensure fidelity in the clinical meaning underneath their data before iterating on the next version of their AI tools.

Photo: claudenakagawa, Getty Images Joseph Zabinski Joseph Zabinski, PhD, MEM, joined IMO Health in 2025 as Senior Vice President of Product Management. He leads the Life Sciences product portfolio and go-to-market strategy. Dr. Zabinski is a published authority and thought leader in AI-driven personalization and outcome prediction using healthcare data. Previously, he oversaw the AI & Personalized Medicine business unit at OM1, a real-world evidence and technology provider. Dr. Zabinski began his career as a consultant at McKinsey, advising pharmaceutical clients on AI strategy and implementation. He holds a PhD from UNC Chapel Hill’s Gillings School of Global Public Health, where his research focused on Bayesian graphical modeling applied to large healthcare datasets. He also earned an MEM in engineering management from Dartmouth, a BS in physics and German studies from Boston College, and a Fulbright Fellowship to Austria.

This post appears through the MedCity Influencers program. Anyone can publish their perspective on business and innovation in healthcare on MedCity News through MedCity Influencers.

Click here to find out how

原始信源MedCity News