Abstract
Identifying causal sentences from nuclear incident reports is essential for advancing nuclear safety research and applications. Nonetheless, accurately locating and labeling causal sentences in text data is challenging, and might benefit from the usage of automated techniques. In this paper, we introduce LERCause, a labeled dataset combined with labeling methods meant to serve as a foundation for the classification of causal sentences in the domain of nuclear safety. We used three BERT models (BERT, BioBERT, and SciBERT) to 10,608 annotated sentences from the Licensee Event Report (LER) corpus for predicting sentence labels (Causal vs. non-Causal). We also used a keyword-based heuristic strategy, three standard machine learning methods (Logistic Regression, Gradient Boosting, and Support Vector Machine), and a deep learning approach (Convolutional Neural Network; CNN) for comparison. We found that the BERT-centric models outperformed all other tested models in terms of all evaluation metrics (accuracy, precision, recall, and F1 score). BioBERT resulted in the highest overall F1 score of 94.49% from the ten-fold cross-validation. Our dataset and coding framework can provide a robust baseline for assessing and comparing new causal sentences extraction techniques. As far as we know, our research breaks new ground by leveraging BERT-centric models for causal sentence classification in the nuclear safety domain and by openly distributing labeled data and code to enable reproducibility in subsequent research.
Publisher
Public Library of Science (PLoS)
Reference65 articles.
1. Data-theoretic approach for socio-technical risk analysis: Text mining licensee event reports of US nuclear power plants;J Pence;Safety science,2020
2. Zhao Y, Diao X, Smidts C. Preliminary Study of Automated Analysis of Nuclear Power Plant Event Reports Based on Natural Language Processing Techniques. Proceedings of the Probabilistic Safety Assessment and Management PSAM. 2018 Sep 16;14.
3. Pence J, Mohaghegh Z, Ostroff C, Dang V, Kee E, Hubenak R, et al. Quantifying organizational factors in human reliability analysis using the big data-theoretic algorithm. International Topical Meeting on Probabilistic Safety Assessment and Analysis, PSA 2015. American Nuclear Society; 2015 Apr. p. 650–9.
4. NUREG C. Licensee Event Report (LER).1989. https://www.nrc.gov/reading-rm/doc-collections/cfr/part050/part050-0073.html
5. Automated Identification of Causal Relationships in Nuclear Power Plant Event Reports;Y Zhao;Nuclear Technology,2019