PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text-Reference-Cited by-同舟云学术

PASCAL: a pseudo cascade learning framework for breast cancer treatment entity normalization in Chinese clinical text

Published:2020-08-28 Issue:1 Volume:20 Page:
ISSN:1472-6947
Container-title:BMC Medical Informatics and Decision Making
language:en
Short-container-title:BMC Med Inform Decis Mak

Author:

An Yang,Wang Jianlin,Zhang Liang^ORCID,Zhao Hanyu,Gao Zhan,Huang Haitao,Du Zhenguang,Jiao Zengtao,Yan Jun,Wei Xiaopeng,Jin Bo

Abstract

Abstract Backgrounds Knowledge discovery from breast cancer treatment records has promoted downstream clinical studies such as careflow mining and therapy analysis. However, the clinical treatment text from electronic health data might be recorded by different doctors under their hospital guidelines, making the final data rich in author- and domain-specific idiosyncrasies. Therefore, breast cancer treatment entity normalization becomes an essential task for the above downstream clinical studies. The latest studies have demonstrated the superiority of deep learning methods in named entity normalization tasks. Fundamentally, most existing approaches adopt pipeline implementations that treat it as an independent process after named entity recognition, which can propagate errors to later tasks. In addition, despite its importance in clinical and translational research, few studies directly deal with the normalization task in Chinese clinical text due to the complexity of composition forms. Methods To address these issues, we propose PASCAL, an end-to-end and accurate framework for breast cancer treatment entity normalization (TEN). PASCAL leverages a gated convolutional neural network to obtain a representation vector that can capture contextual features and long-term dependencies. Additionally, it treats treatment entity recognition (TER) as an auxiliary task that can provide meaningful information to the primary TEN task and as a particular regularization to further optimize the shared parameters. Finally, by concatenating the context-aware vector and probabilistic distribution vector from TEN, we utilize the conditional random field layer (CRF) to model the normalization sequence and predict the TEN sequential results. Results To evaluate the effectiveness of the proposed framework, we employ the three latest sequential models as baselines and build the model in single- and multitask on a real-world database. Experimental results show that our method achieves better accuracy and efficiency than state-of-the-art approaches. Conclusions The effectiveness and efficiency of the presented pseudo cascade learning framework were validated for breast cancer treatment normalization in clinical text. We believe the predominant performance lies in its ability to extract valuable information from unstructured text data, which will significantly contribute to downstream tasks, such as treatment recommendations, breast cancer staging and careflow mining.

Publisher

Springer Science and Business Media LLC

Subject

Health Informatics,Health Policy,Computer Science Applications

Link

https://link.springer.com/content/pdf/10.1186/s12911-020-01216-9.pdf

Reference34 articles.

1. Marklund L, Hammarstedt L. Impact of hpv in oropharyngeal cancer. J Oncol. 2011; 2011(1687-8450):509036. https://doi.org/10.1155/2011/509036.

2. What Is Breast Cancer?https://www.imaginis.com/general-information-on-breast-cancer/what-is-breast-cancer-2. Accessed 11 June 2008.

3. Dagliati A, Sacchi L, Zambelli A, Tibollo V, Pavesi L, Holmes JH, Bellazzi R. Temporal electronic phenotyping by mining careflows of breast cancer patients. J Biomed Inform; 66:136–47. https://doi.org/10.1016/j.jbi.2016.12.012.

4. Yadav R, Khan Z, Saxena H. Chemotherapy prediction of cancer patient by using data mining techniques. Int J Comput Appl. 2014; 76(10):28–31. https://doi.org/10.5120/13285-0747.

5. Wang XH, Zheng B, Good WF, King JL, Chang Y-H. Computer-assisted diagnosis of breast cancer using a data-driven bayesian belief network. Int J Med Inform; 54(2):115–26. https://doi.org/10.1016/S1386-5056(98)00174-9.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Applications of different machine learning approaches in prediction of breast cancer diagnosis delay;Frontiers in Oncology;2023-02-16

2. Chemical Named Entity Recognition for Ovarian Cancer’s Drug Discovery;2022 International Conference on Decision Aid Sciences and Applications (DASA);2022-03-23