Automated detection of over- and under-dispersion in baseline tables in randomised controlled trials-Reference-Cited by-同舟云学术

Automated detection of over- and under-dispersion in baseline tables in randomised controlled trials

Published:2023-05-30 Issue: Volume:11 Page:783
ISSN:2046-1402
Container-title:F1000Research
language:en
Short-container-title:F1000Res

Author:

Barnett Adrian^ORCID

Abstract

Background: Papers describing the results of a randomised trial should include a baseline table that compares the characteristics of randomised groups. Researchers who fraudulently generate trials often unwittingly create baseline tables that are implausibly similar (under-dispersed) or have large differences between groups (over-dispersed). I aimed to create an automated algorithm to screen for under- and over-dispersion in the baseline tables of randomised trials. Methods: Using a cross-sectional study I examined 2,245 randomised controlled trials published in health and medical journals on PubMed Central. I estimated the probability that a trial's baseline summary statistics were under- or over-dispersed using a Bayesian model that examined the distribution of t-statistics for the between-group differences, and compared this with an expected distribution without dispersion. I used a simulation study to test the ability of the model to find under- or over-dispersion and compared its performance with an existing test of dispersion based on a uniform test of p-values. My model combined categorical and continuous summary statistics, whereas the uniform test used only continuous statistics. Results: The algorithm had a relatively good accuracy for extracting the data from baseline tables, matching well on the size of the tables and sample size. Using t-statistics in the Bayesian model out-performed the uniform test of p-values, which had many false positives for skewed, categorical and rounded data that were not under- or over-dispersed. For trials published on PubMed Central, some tables appeared under- or over-dispersed because they had an atypical presentation or had reporting errors. Some trials flagged as under-dispersed had groups with strikingly similar summary statistics. Conclusions: Automated screening for fraud of all submitted trials is challenging due to the widely varying presentation of baseline tables. The Bayesian model could be useful in targeted checks of suspected trials or authors.

Funder

National Health and Medical Research Council

Publisher

F1000 Research Ltd

Subject

General Pharmacology, Toxicology and Pharmaceutics,General Immunology and Microbiology,General Biochemistry, Genetics and Molecular Biology,General Medicine

Link

https://f1000research.com/articles/11-783/v2/pdf

Reference52 articles.

1. Subgroup analysis, covariate adjustment and baseline comparisons in clinical trial reporting: current practiceand problems.;S Pocock;Stat. Med.,2002

2. CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials.;K Schulz;BMJ.,2010

3. Just post it.;U Simonsohn;Psychol. Sci.,2013

4. How a data detective exposed suspicious medical trials.;D Adam;Nature.,2019

5. False individual patient data and zombie randomised controlled trials submitted to Anaesthesia.;J Carlisle;Anaesthesia.,2020

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A survey of experts to identify methods to detect problematic studies: Stage 1 of the INSPECT-SR Project;2024-03-19