Evaluation of State-of-the-Art Paraphrase Identification and Its Application to Automatic Plagiarism Detection-Reference-Cited by-同舟云学术

Evaluation of State-of-the-Art Paraphrase Identification and Its Application to Automatic Plagiarism Detection

Published:2019-08-22 Issue:04 Volume:34 Page:2053004
ISSN:0218-0014
Container-title:International Journal of Pattern Recognition and Artificial Intelligence
language:en
Short-container-title:Int. J. Patt. Recogn. Artif. Intell.

Author:

Altheneyan Alaa¹,Menai Mohamed El Bachir¹^ORCID

Affiliation:

1. Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, P. O. Box 89638, Saudi Arabia

Abstract

Paraphrase identification is a natural language processing (NLP) problem that involves the determination of whether two text segments have the same meaning. Various NLP applications rely on a solution to this problem, including automatic plagiarism detection, text summarization, machine translation (MT), and question answering. The methods for identifying paraphrases found in the literature fall into two main classes: similarity-based methods and classification methods. This paper presents a critical study and an evaluation of existing methods for paraphrase identification and its application to automatic plagiarism detection. It presents the classes of paraphrase phenomena, the main methods, and the sets of features used by each particular method. All the methods and features used are discussed and enumerated in a table for easy comparison. Their performances on benchmark corpora are also discussed and compared via tables. Automatic plagiarism detection is presented as an application of paraphrase identification. The performances on benchmark corpora of existing plagiarism detection systems able to detect paraphrases are compared and discussed. The main outcome of this study is the identification of word overlap, structural representations, and MT measures as feature subsets that lead to the best performance results for support vector machines in both paraphrase identification and plagiarism detection on corpora. The performance results achieved by deep learning techniques highlight that these techniques are the most promising research direction in this field.

Publisher

World Scientific Pub Co Pte Lt

Subject

Artificial Intelligence,Computer Vision and Pattern Recognition,Software

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0218001420530043

Reference85 articles.

1. Paraphrase identification and semantic text similarity analysis in Arabic news tweets using lexical, syntactic, and semantic features

2. A bit-string longest-common-subsequence algorithm

3. Understanding Plagiarism Linguistic Patterns, Textual Features, and Detection Methods

4. A Survey of Paraphrasing and Textual Entailment Methods

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Building trust with the ethical affordances of education technologies: A sociotechnical systems perspective;Putting AI in the Critical Loop;2024

2. Comparison study of unsupervised paraphrase detection: Deep learning—The key for semantic similarity detection;Expert Systems;2023-06-22

3. Towards diverse and contextually anchored paraphrase modeling: A dataset and baselines for Finnish;Natural Language Engineering;2023-03-16

4. Evaluation of Different Plagiarism Detection Methods: A Fuzzy MCDM Perspective;Applied Sciences;2022-04-30

5. Assisting academics to identify computer generated writing;European Journal of Engineering Education;2022-03-02