Domain-specific LLM Development and Evaluation – A Case-study for Prostate Cancer-Reference-Cited by-同舟云学术

Domain-specific LLM Development and Evaluation – A Case-study for Prostate Cancer

Published:2024-03-16 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Tariq Amara,Luo Man,Urooj Aisha,Das Avisha,Jeong Jiwoong,Trivedi Shubham,Patel Bhavik,Banerjee Imon

Abstract

AbstractIn this work, we present our strategy for developing domain-specific large language models which cover the vocabulary of the target domain and train on reliable sources of clinical information. Prostate cancer was chosen as a use-case for this study. We collected more than 1.8 million clinical notes and radiology and pathology reports for 15341 patients treated for prostate cancer in Mayo Clinic across three sites and outpatient clinics. In addition to domain-specific training data, we built domain-specific tokenizers and devised knowledge-guided training strategies for LLM development. During the self-supervised training, LLM was forced to predict domain-specific information by marking clinical terms using UMLS parser. We evaluated the model for downstream tasks of clinical information prediction and question answering using quantitative and user evaluation study to measure the accuracy, reliability and information completeness. We compared the domain-specific model against similarly sized general purpose model GPT-2 and a three-times larger domain specialized model. i.e., BioGPT. Our model outperformed GPT-2 on both tasks by a wide margin. Our model was also able to outperform BioGPT on clinical information prediction tasks and showed some advantages over BioGPT in question-answering tasks.

Publisher

Cold Spring Harbor Laboratory

Reference26 articles.

1. Trends in Prostate Cancer in the United States

2. Annual report to the nation on the status of cancer, part I: National cancer statistics

3. Predictors of the Use of a Mental Health–Focused eHealth System in Patients With Breast and Prostate Cancer: Bayesian Structural Equation Modeling Analysis of a Prospective Study;JMIR cancer,2023

4. Risks of alcohol and drug use disorders in prostate cancer survivors: a national cohort study;JNCI Cancer Spectrum,2023

5. Objective data reveals gender preferences for patients’ primary care physician;Journal of Primary Care & Community Health,2020