A Benchmark Dataset to Distinguish Human-Written and Machine-Generated Scientific Papers-Reference-Cited by-同舟云学术

A Benchmark Dataset to Distinguish Human-Written and Machine-Generated Scientific Papers

Published:2023-09-26 Issue:10 Volume:14 Page:522
ISSN:2078-2489
Container-title:Information
language:en
Short-container-title:Information

Author:

Abdalla Mohamed Hesham Ibrahim¹,Malberg Simon¹^ORCID,Dementieva Daryna¹^ORCID,Mosca Edoardo¹^ORCID,Groh Georg¹

Affiliation:

1. School of Computation, Information and Technology, Technical University of Munich, 80333 Munich, Germany

Abstract

As generative NLP can now produce content nearly indistinguishable from human writing, it is becoming difficult to identify genuine research contributions in academic writing and scientific publications. Moreover, information in machine-generated text can be factually wrong or even entirely fabricated. In this work, we introduce a novel benchmark dataset containing human-written and machine-generated scientific papers from SCIgen, GPT-2, GPT-3, ChatGPT, and Galactica, as well as papers co-created by humans and ChatGPT. We also experiment with several types of classifiers—linguistic-based and transformer-based—for detecting the authorship of scientific text. A strong focus is put on generalization capabilities and explainability to highlight the strengths and weaknesses of these detectors. Our work makes an important step towards creating more robust methods for distinguishing between human-written and machine-generated scientific papers, ultimately ensuring the integrity of scientific literature.

Funder

Federal Ministry of Education and Research

Publisher

MDPI AG

Subject

Information Systems

Link

https://www.mdpi.com/2078-2489/14/10/522/pdf

Reference74 articles.

1. Language models are few-shot learners;Brown;Adv. Neural Inf. Process. Syst.,2020

2. Scao, T.L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A.S., Yvon, F., and Gallé, M. (2022). Bloom: A 176b-parameter open-access multilingual language model. arXiv.

3. OpenAI (2023). GPT-4 Technical Report. arXiv.

4. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019). Language Models are Unsupervised Multitask Learners. OpenAI Blog, Available online: https://insightcivic.s3.us-east-1.amazonaws.com/language-models.pdf.

5. Keskar, N.S., McCann, B., Varshney, L.R., Xiong, C., and Socher, R. (2019). CTRL: A Conditional Transformer Language Model for Controllable Generation. arXiv.

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Conversational and generative artificial intelligence and human–chatbot interaction in education and research;International Transactions in Operational Research;2024-07-31

2. Preface to the Special Issue on Computational Linguistics and Natural Language Processing;Information;2024-05-15

3. The Explainability of Transformers: Current Status and Directions;Computers;2024-04-04

4. FELIX: Automatic and Interpretable Feature Engineering Using LLMs;Lecture Notes in Computer Science;2024