A Self-Supervised Representation Learning of Sentence Structure for Authorship Attribution-Reference-Cited by-同舟云学术

A Self-Supervised Representation Learning of Sentence Structure for Authorship Attribution

Published:2022-01-08 Issue:4 Volume:16 Page:1-16
ISSN:1556-4681
Container-title:ACM Transactions on Knowledge Discovery from Data
language:en
Short-container-title:ACM Trans. Knowl. Discov. Data

Author:

Jafariakinabad Fereshteh¹^ORCID,Hua Kien A.¹

Affiliation:

1. University of Central Florida, Orlando, Florida, FL

Abstract

The syntactic structure of sentences in a document substantially informs about its authorial writing style. Sentence representation learning has been widely explored in recent years and it has been shown that it improves the generalization of different downstream tasks across many domains. Even though utilizing probing methods in several studies suggests that these learned contextual representations implicitly encode some amount of syntax, explicit syntactic information further improves the performance of deep neural models in the domain of authorship attribution. These observations have motivated us to investigate the explicit representation learning of syntactic structure of sentences. In this article, we propose a self-supervised framework for learning structural representations of sentences. The self-supervised network contains two components; a lexical sub-network and a syntactic sub-network which take the sequence of words and their corresponding structural labels as the input, respectively. Due to the n -to-1 mapping of words to their structural labels, each word will be embedded into a vector representation which mainly carries structural information. We evaluate the learned structural representations of sentences using different probing tasks, and subsequently utilize them in the authorship attribution task. Our experimental results indicate that the structural embeddings significantly improve the classification tasks when concatenated with the existing pre-trained word embeddings.

Funder

Crystal Photonics Inc

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3491203

Reference57 articles.

1. Detecting Hoaxes, Frauds, and Deception in Writing Style Online

2. Shlomo Argamon-Engelson, Moshe Koppel, and Galit Avneri. 1998. Style-based text categorization: What newspaper am I reading. In Proceedings of the AAAI Workshop on Text Categorization. 1–4.

3. Authorship clustering using multi-headed recurrent neural networks;Bagnall Douglas;arXiv:1608.04485,2016

4. Neural machine translation by jointly learning to align and translate;Bahdanau Dzmitry;arXiv:1409.0473,2014

5. Generating Sentences from Disentangled Syntactic and Semantic Spaces

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Understanding writing style in social media with a supervised contrastively pre-trained transformer;Knowledge-Based Systems;2024-07

2. Knowledge Graph-Based Hierarchical Text Semantic Representation;International Journal of Intelligent Systems;2024-01-12

3. Determination of the Features of the Author’s Style of A.S. Pushkin’s Poems by Machine Learning Methods;Applied Sciences;2022-02-06