Bat4RCT: A suite of benchmark data and baseline methods for text classification of randomized controlled trials-Reference-Cited by-同舟云学术

Bat4RCT: A suite of benchmark data and baseline methods for text classification of randomized controlled trials

Published:2023-03-24 Issue:3 Volume:18 Page:e0283342
ISSN:1932-6203
Container-title:PLOS ONE
language:en
Short-container-title:PLoS ONE

Author:

Kim Jenna^ORCID,Kim Jinmo,Lee Aejin,Kim Jinseok^ORCID

Abstract

Randomized controlled trials (RCTs) play a major role in aiding biomedical research and practices. To inform this research, the demand for highly accurate retrieval of scientific articles on RCT research has grown in recent decades. However, correctly identifying all published RCTs in a given domain is a non-trivial task, which has motivated computer scientists to develop methods for identifying papers involving RCTs. Although existing studies have provided invaluable insights into how RCT tags can be predicted for biomedicine research articles, they used datasets from different sources in varying sizes and timeframes and their models and findings cannot be compared across studies. In addition, as datasets and code are rarely shared, researchers who conduct RCT classification have to write code from scratch, reinventing the wheel. In this paper, we present Bat4RCT, a suite of data and an integrated method to serve as a strong baseline for RCT classification, which includes the use of BERT-based models in comparison with conventional machine learning techniques. To validate our approach, all models are applied on 500,000 paper records in MEDLINE. The BERT-based models showed consistently higher recall scores than conventional machine learning and CNN models while producing slightly better or similar precision scores. The best performance was achieved by the BioBERT model when trained on both title and abstract texts, with the F1 score of 90.85%. This infrastructure of dataset and code will provide a competitive baseline for the evaluation and comparison of new methods and the convenience of future benchmarking. To our best knowledge, our study is the first work to apply BERT-based language modeling techniques to RCT classification tasks and to share dataset and code in order to promote reproducibility and improvement in text classification in biomedicine research.

Publisher

Public Library of Science (PLoS)

Subject

Multidisciplinary

Reference25 articles.

1. Randomized controlled trials.;HO Stolberg;Am J Roentgenol,2004

2. Automated confidence ranked classification of randomized controlled trial articles: an aid to evidence-based medicine.;AM Cohen;J Am Med Inform Assoc,2015

3. Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews;J Thomas;J Clin Epidemiol,2021

4. Identifying reports of randomized controlled trials (RCTs) via a hybrid machine learning and crowdsourcing approach.;BC Wallace;J Am Med Inform Assoc,2017

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. LERCause: Deep learning approaches for causal sentence identification from nuclear safety reports;PLOS ONE;2024-08-22

2. A Pipeline for the Automatic Identification of Randomized Controlled Oncology Trials and Assignment of Tumor Entities Using Natural Language Processing;2024-07-03

3. Clinical Text Analysis with Natural Language Processing: A BERT-based Approach;2024 International Conference on Communication, Computer Sciences and Engineering (IC3SE);2024-05-09