Improving Clinical Note Generation from Complex Doctor-Patient Conversation

LLM SFT NLP CLFDTG
2024年08月26日
撰写临床笔记和记录医学检查是医疗保健专业人员的关键任务,是患者护理文档的重要组成部分。然而,手动撰写这些笔记是耗时的,可能会影响临床医生与患者直接互动和完成其他任务的时间。因此,在AI健康领域中,自动临床笔记生成系统的开发已成为一个具有临床意义的研究领域。在本文中,我们提出了三个关键贡献,利用大型语言模型(LLMs)进行临床笔记生成的领域。首先,我们介绍了CliniKnote,这是一个包含1,200个复杂医生-患者对话及其完整临床笔记的全面数据集。该数据集由医学专家在现代神经网络的帮助下创建和策划,为临床笔记生成任务的模型训练和评估提供了有价值的资源。其次,我们提出了K-SOAP(关键词、主观、客观、评估和计划)笔记格式,它通过在顶部添加一个关键词部分,增强了传统SOAP(主观、客观、评估和计划)笔记,从而允许快速识别关键信息。第三,我们开发了一个自动流水线,从医生-患者对话中生成K-SOAP笔记,并使用各种指标对各种现代LLMs进行基准测试。我们的结果表明,与标准LLM微调方法相比,效率和性能都有显着提高。
Writing clinical notes and documenting medical exams is a critical task for healthcare professionals, serving as a vital component of patient care documentation. However, manually writing these notes is time-consuming and can impact the amount of time clinicians can spend on direct patient interaction and other tasks. Consequently, the development of automated clinical note generation systems has emerged as a clinically meaningful area of research within AI for health. In this paper, we present three key contributions to the field of clinical note generation using large language models (LLMs). First, we introduce CliniKnote, a comprehensive dataset consisting of 1,200 complex doctor-patient conversations paired with their full clinical notes. This dataset, created and curated by medical experts with the help of modern neural networks, provides a valuable resource for training and evaluating models in clinical note generation tasks. Second, we propose the K-SOAP (Keyword, Subjective, Objective, Assessment, and Plan) note format, which enhances traditional SOAP~\cite{podder2023soap} (Subjective, Objective, Assessment, and Plan) notes by adding a keyword section at the top, allowing for quick identification of essential information. Third, we develop an automatic pipeline to generate K-SOAP notes from doctor-patient conversations and benchmark various modern LLMs using various metrics. Our results demonstrate significant improvements in efficiency and performance compared to standard LLM finetuning methods.
许愿