Improving the Reliability of Peer Review Without a Gold Standard-Reference-Cited by-同舟云学术

Improving the Reliability of Peer Review Without a Gold Standard

Published:2024-02-05 Issue:2 Volume:37 Page:489-503
ISSN:2948-2933
Container-title:Journal of Imaging Informatics in Medicine
language:en
Short-container-title:J Digit Imaging. Inform. med.

Author:

Äijö Tarmo^ORCID,Elgort Daniel,Becker Murray,Herzog Richard,Brown Richard K. J.,Odry Benjamin L.,Vianu Ron

Abstract

AbstractPeer review plays a crucial role in accreditation and credentialing processes as it can identify outliers and foster a peer learning approach, facilitating error analysis and knowledge sharing. However, traditional peer review methods may fall short in effectively addressing the interpretive variability among reviewing and primary reading radiologists, hindering scalability and effectiveness. Reducing this variability is key to enhancing the reliability of results and instilling confidence in the review process. In this paper, we propose a novel statistical approach called “Bayesian Inter-Reviewer Agreement Rate” (BIRAR) that integrates radiologist variability. By doing so, BIRAR aims to enhance the accuracy and consistency of peer review assessments, providing physicians involved in quality improvement and peer learning programs with valuable and reliable insights. A computer simulation was designed to assign predefined interpretive error rates to hypothetical interpreting and peer-reviewing radiologists. The Monte Carlo simulation then sampled (100 samples per experiment) the data that would be generated by peer reviews. The performances of BIRAR and four other peer review methods for measuring interpretive error rates were then evaluated, including a method that uses a gold standard diagnosis. Application of the BIRAR method resulted in 93% and 79% higher relative accuracy and 43% and 66% lower relative variability, compared to “Single/Standard” and “Majority Panel” peer review methods, respectively. Accuracy was defined by the median difference of Monte Carlo simulations between measured and pre-defined “actual” interpretive error rates. Variability was defined by the 95% CI around the median difference of Monte Carlo simulations between measured and pre-defined “actual” interpretive error rates. BIRAR is a practical and scalable peer review method that produces more accurate and less variable assessments of interpretive quality by accounting for variability within the group’s radiologists, implicitly applying a standard derived from the level of consensus within the group across various types of interpretive findings.

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s10278-024-00971-9.pdf

Reference23 articles.

1. Sarwar A, Boland G, Monks A, Kruskal JB. Metrics for radiologists in the era of value-based health care delivery. Radiographics. 2015;35(3). https://doi.org/10.1148/rg.2015140221

2. Brady A, Brink J, Slavotinek J. Radiology and Value-Based Health Care. JAMA - Journal of the American Medical Association. 2020;324(13). https://doi.org/10.1001/jama.2020.14930

3. Ericsson KA. INVITED ADDRESS Deliberate Practice and the Acquisition and Maintenance of Expert Performance in Medicine and Related Domains.; 2003. http://journals.lww.com/academicmedicine

4. Karsh BT, Holden RJ, Alper SJ, Or CKL. A human factors engineering paradigm for patient safety: Designing to support the performance of the healthcare professional. Qual Saf Health Care. 2006;15(SUPPL. 1). https://doi.org/10.1136/qshc.2005.015974

5. Bender LC, Linnau KF, Meier EN, Anzai Y, Gunn ML. Interrater agreement in the evaluation of discrepant imaging findings with the Radpeer system. American Journal of Roentgenology. 2012;199(6). https://doi.org/10.2214/AJR.12.8972