A large dataset of scientific text reuse in Open-Access publications-Reference-Cited by-同舟云学术

A large dataset of scientific text reuse in Open-Access publications

Published:2023-01-26 Issue:1 Volume:10 Page:
ISSN:2052-4463
Container-title:Scientific Data
language:en
Short-container-title:Sci Data

Author:

Gienapp Lukas^ORCID,Kircheis Wolfgang,Sievers Bjarne,Stein Benno,Potthast Martin

Abstract

AbstractWe present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications. It contains 91 million cases of reused text passages found in 4.2 million unique open-access publications. Cases range from overlap of as few as eight words to near-duplicate publications and include a variety of reuse types, ranging from boilerplate text to verbatim copying to quotations and paraphrases. Featuring a high coverage of scientific disciplines and varieties of reuse, as well as comprehensive metadata to contextualize each case, our dataset addresses the most salient shortcomings of previous ones on scientific writing. The Webis-STEREO-21 does not indicate if a reuse case is legitimate or not, as its focus is on the general study of text reuse in science, which is legitimate in the vast majority of cases. It allows for tackling a wide range of research questions from different scientific backgrounds, facilitating both qualitative and quantitative analysis of the phenomenon as well as a first-time grounding on the base rate of text reuse in scientific publications.

Funder

Bundesministerium für Bildung und Forschung

Publisher

Springer Science and Business Media LLC

Subject

Library and Information Sciences,Statistics, Probability and Uncertainty,Computer Science Applications,Education,Information Systems,Statistics and Probability

Link

https://www.nature.com/articles/s41597-022-01908-z.pdf

Reference42 articles.

1. Sun, Y.-C. & Yang, F.-Y. Uncovering published authors’ text-borrowing practices: Paraphrasing strategies, sources, and self-plagiarism. Journal of English for Academic Purposes 20, 224–236, https://doi.org/10.1016/j.jeap.2015.05.003 (2015).

2. Anson, I. G. & Moskovitz, C. Text recycling in stem: a text-analytic study of recently published research articles. Accountability in Research 28, 349–371 (2020).

3. Moskovitz, C. Self-plagiarism, text recycling, and science education. BioScience 66, 5–6, https://doi.org/10.1093/biosci/biv160 (2015).

4. Hall, S., Moskovitz, C. & Pemberton, M. A. Attitudes toward text recycling in academic writing across disciplines. Accountability in Research 25, 142–169, https://doi.org/10.1080/08989621.2018.1434622 (2018).

5. Bird, S. J. Self-plagiarism and dual and redundant publications: what is the problem? Science and engineering ethics 8, 543–544 (2002).

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Analyzing Mathematical Content for Plagiarism and Recommendations;Lecture Notes in Computer Science;2024

2. Mutual Influence in Citation and Cooperation Patterns;IEEE Transactions on Computational Social Systems;2023