INCEPT: A Framework for Duplicate Posts Classification with Combined Text Representations-Reference-Cited by-同舟云学术

INCEPT: A Framework for Duplicate Posts Classification with Combined Text Representations

Published:2024-08-16 Issue:3 Volume:18 Page:1-24
ISSN:1559-1131
Container-title:ACM Transactions on the Web
language:en
Short-container-title:ACM Trans. Web

Author:

Skenderi Erjon¹^ORCID,Huhtamäki Jukka²^ORCID,Laaksonen Salla-Maaria³^ORCID,Stefanidis Kostas²^ORCID

Affiliation:

1. Tampere University, Tampere, Finland and University of Helsinki, Helsinki, Finland

2. Tampere University, Tampere, Finland

3. University of Helsinki, Helsinki, Finland

Abstract

Dealing with many of the problems related to the quality of textual content online involves identifying similar content. Algorithmic solutions for duplicate content classification typically rely on text vector representation, which maps textual information into a set of features. Ideally, this representation would capture all aspects of the underlying text, including length, word frequencies, syntax, and semantics. While recent advancements in text representation have led to improved performance, a comprehensive approach that explicitly incorporates all text features has not yet been proposed. In this study, we present the INCEPT framework that utilizes multiple representation methods to detect duplicate text pairs, taking advantage of their individual strengths. The core of our approach involves using a stacking ensemble of pairwise vector distance measurements that are computed from multiple text representation methods. A stacking classifier then utilizes these distance scores as input and learns to identify duplicate posts. We assess the proposed framework’s effectiveness in identifying duplicate posts in an online Question and Answer platform. By combining several text representation methods, INCEPT performs well in the duplicate posts classification task. Our experiments demonstrate that specific framework configurations outperform the accuracy scores obtained from individual text representation methods. Therefore, we also infer that no single text representation method can independently capture a text’s features.

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3677322

Reference55 articles.

1. Proceedings of the 13th International Conference on Mining Software Repositories

2. “The Enemy Among Us”

3. 10.1162/153244303322533223

4. Enriching Word Vectors with Subword Information