Inter-reviewer reliability of human literature reviewing and implications for the introduction of machine-assisted systematic reviews: a mixed-methods review-Reference-Cited by-同舟云学术

Inter-reviewer reliability of human literature reviewing and implications for the introduction of machine-assisted systematic reviews: a mixed-methods review

Published:2024-03 Issue:3 Volume:14 Page:e076912
ISSN:2044-6055
Container-title:BMJ Open
language:en
Short-container-title:BMJ Open

Author:

Hanegraaf Piet,Wondimu Abrham,Mosselman Jacob Jan,de Jong Rutger,Abogunrin Seye,Queiros Luisa,Lane Marie,Postma Maarten J,Boersma Cornelis,van der Schans Jurjen^ORCID

Abstract

ObjectivesOur main objective is to assess the inter-reviewer reliability (IRR) reported in published systematic literature reviews (SLRs). Our secondary objective is to determine the expected IRR by authors of SLRs for both human and machine-assisted reviews.MethodsWe performed a review of SLRs of randomised controlled trials using the PubMed and Embase databases. Data were extracted on IRR by means of Cohen’s kappa score of abstract/title screening, full-text screening and data extraction in combination with review team size, items screened and the quality of the review was assessed with the A MeaSurement Tool to Assess systematic Reviews 2. In addition, we performed a survey of authors of SLRs on their expectations of machine learning automation and human performed IRR in SLRs.ResultsAfter removal of duplicates, 836 articles were screened for abstract, and 413 were screened full text. In total, 45 eligible articles were included. The average Cohen’s kappa score reported was 0.82 (SD=0.11, n=12) for abstract screening, 0.77 (SD=0.18, n=14) for full-text screening, 0.86 (SD=0.07, n=15) for the whole screening process and 0.88 (SD=0.08, n=16) for data extraction. No association was observed between the IRR reported and review team size, items screened and quality of the SLR. The survey (n=37) showed overlapping expected Cohen’s kappa values ranging between approximately 0.6–0.9 for either human or machine learning-assisted SLRs. No trend was observed between reviewer experience and expected IRR. Authors expect a higher-than-average IRR for machine learning-assisted SLR compared with human based SLR in both screening and data extraction.ConclusionCurrently, it is not common to report on IRR in the scientific literature for either human and machine learning-assisted SLRs. This mixed-methods review gives first guidance on the human IRR benchmark, which could be used as a minimal threshold for IRR in machine learning-assisted SLRs.PROSPERO registration numberCRD42023386706.

Funder

F. Hoffmann-La Roche

Publisher

BMJ

Reference24 articles.

1. Evidence based medicine: what it is and what it isn't

2. Systematic Research Synthesis to Inform Policy, Practice and Democratic Debate

3. How Quickly Do Systematic Reviews Go Out of Date? A Survival Analysis

4. Living systematic review: 1. Introduction—the why, what, when, and how

5. Machine learning computational tools to assist the performance of systematic reviews: a mapping review;Cierco Jimenez;BMC Med Res Methodol,2022