A Vision-language Framework for Comparative Reasoning in Radiology

Multimodal Intelligence Vision-Language Pre-training VQA CV LDCAD
2026年06月04日
医学影像人工智能在单幅图像的独立解读任务中已展现出卓越性能,但其与放射科临床实践仍存在显著脱节——因为实际诊断和随访工作高度依赖于对既往检查结果及类似参考病例的跨时间、跨案例比对分析。本文将放射科比对任务建模为一种“实体感知型跨图像推理”问题,并提出一个统一框架,同步支持参考病例检索与纵向对比性影像解读。我们构建了MedReCo-DB——一个大规模、面向临床比对任务的影像资源库,该库源自常规临床实践中积累的影像-报告配对数据,涵盖来自8家医疗机构、4个国家、7种影像模态的逾69万张影像,涉及超16万名患者。所有报告均被系统性地结构化解析为解剖结构、异常征象与病理状态三类语义实体,从而为实体条件约束下的精准检索与对比式视觉问答任务提供高质量监督信号。基于该资源库,我们研发了MedReCo——一种具备实体感知能力的视觉编码器,可实现临床意义高度相似病例的可控式检索;同时开发了MedReCo-VLM——一种视觉-语言融合模型,专用于生成式地解读病灶随时间推移的动态变化。在内部验证、外部验证及跨中心验证三类评估中,MedReCo在全部12项内部检索任务中均取得最高的Recall@1指标;在外源性检索任务中,平均提升幅度达6.0个百分点;尤其在临床易混淆的鉴别诊断组别中,其表现持续显著优于最强基线模型。MedReCo-VLM则在全部对比生成类评估任务中均斩获最佳性能:在胸部X线片的纵向随访任务中,准确率提升14.5–46.5个百分点;在CT影像中亦提升13.0–27.9个百分点。上述结果表明,实体感知型的对比推理能力完全可从海量常规临床数据中规模化习得,有望为医学影像人工智能构建更契合真实临床需求的技术基础。
Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across prior studies and analogous reference cases. Here we formulate radiological comparison as an entity-aware cross-image reasoning problem and introduce a framework that supports both reference-case retrieval and temporal comparative interpretation. We construct MedReCo-DB, a large-scale comparative imaging resource derived from routine image-report pairs, comprising more than 690,000 images from over 160,000 patients across eight institutions, four countries and seven imaging modalities. Reports are decomposed into anatomical structures, abnormal findings and pathological conditions to provide supervision for entity-conditioned retrieval and comparative visual question answering. Using this resource, we develop MedReCo, an entity-aware visual encoder for controllable retrieval of clinically analogous cases, and MedReCo-VLM, a vision--language extension for generative interpretation of interval change. Across internal, external and cross-center evaluations, MedReCo achieved the highest Recall@1 in all 12 internal retrieval settings and improved external retrieval by a mean of 6.0 percentage points. In clinically confusable differential groups, it consistently outperformed the strongest baselines. MedReCo-VLM achieved the best performance across all comparative generation evaluations and improved longitudinal follow-up accuracy by 14.5-46.5 percentage points on chest radiographs and 13.0-27.9 percentage points on CT. These findings suggest that entity-aware comparative reasoning can be learned from routine clinical data at scale and may provide a more clinically aligned foundation for medical imaging AI.
许愿