Knowledge-enhanced visual-language pre-training on chest radiology images-Reference-Cited by-同舟云学术

Knowledge-enhanced visual-language pre-training on chest radiology images

Published:2023-07-28 Issue:1 Volume:14 Page:
ISSN:2041-1723
Container-title:Nature Communications
language:en
Short-container-title:Nat Commun

Author:

Zhang Xiaoman^ORCID,Wu Chaoyi,Zhang Ya,Xie Weidi^ORCID,Wang Yanfeng^ORCID

Abstract

AbstractWhile multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nature of medical tasks and the high demand for domain knowledge. To address this challenge, we propose an approach called Knowledge-enhanced Auto Diagnosis (KAD) which leverages existing medical domain knowledge to guide vision-language pre-training using paired chest X-rays and radiology reports. We evaluate KAD on four external X-ray datasets and demonstrate that its zero-shot performance is not only comparable to that of fully supervised models but also superior to the average of three expert radiologists for three (out of five) pathologies with statistical significance. Moreover, when few-shot annotation is available, KAD outperforms all existing approaches in fine-tuning settings, demonstrating its potential for application in different clinical scenarios.

Publisher

Springer Science and Business Media LLC

Subject

General Physics and Astronomy,General Biochemistry, Genetics and Molecular Biology,General Chemistry,Multidisciplinary

Link

https://www.nature.com/articles/s41467-023-40260-7.pdf

Reference37 articles.

1. Bommasani, R. et al. On the opportunities and risks of foundation models. Preprint at https://arxiv.org/abs/2108.07258 (2021).

2. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–4186 (ACL, 2019).

3. Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33, 1877–1901 (2020).

4. Radford, A. et al. Learning transferable visual models from natural language supervision. PMLR 139, 8748–8763 (2021).

5. Ma, C., Yang, Y., Wang, Y., Zhang, Y. & Xie, W. Open-vocabulary semantic segmentation with frozen vision-language models. British Machine Vision Conference (2022).

Cited by 16 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Enhancing the vision–language foundation model with key semantic knowledge-emphasized report refinement;Medical Image Analysis;2024-10

2. Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge;Medical Image Analysis;2024-10

3. Enhancing representation in radiography-reports foundation model: a granular alignment algorithm using masked contrastive learning;Nature Communications;2024-09-02

4. Interactive dual-stream contrastive learning for radiology report generation;Journal of Biomedical Informatics;2024-09

5. Graph Artificial Intelligence in Medicine;Annual Review of Biomedical Data Science;2024-08-23