A Comprehensive Study of Learning Approaches for Author Gender Identification-Reference-Cited by-同舟云学术

A Comprehensive Study of Learning Approaches for Author Gender Identification

Published:2022-09-23 Issue:3 Volume:51 Page:429-445
ISSN:2335-884X
Container-title:Information Technology and Control
language:
Short-container-title:ITC

Author:

Dalyan Tuğba,Ayral Hakan,Özdemir Özgür

Abstract

In recent years, author gender identification is an important yet challenging task in the fields of information retrieval and computational linguistics. In this paper, different learning approaches are presented to address the problem of author gender identification for Turkish articles. First, several classification algorithms are applied to the list of representations based on different paradigms: fixed-length vector representations such as Stylometric Features (SF), Bag-of-Words (BoW) and distributed word/document embeddings such as Word2vec, fastText and Doc2vec. Secondly, deep learning architectures, Convolution Neural Network (CNN), Recurrent Neural Network (RNN), special kinds of RNN such as Long-Short Term Memory (LSTM) and Gated Recurrent Unit (GRU), C-RNN, Bidirectional LSTM (bi-LSTM), Bidirectional GRU (bi-GRU), Hierarchical Attention Networks and Multi-head Attention (MHA) are designated and their comparable performances are evaluated. We conducted a variety of experiments and achieved outstanding empirical results. To conclude, ML algorithms with BoW have promising results. fast-Text is also probably suitable between embedding models. This comprehensive study contributes to literature utilizing different learning approaches based on several ways of representations. It is also first important attempt to identify author gender applying SF, fastText and DNN architectures to the Turkish language.

Publisher

Kaunas University of Technology (KTU)

Subject

Electrical and Electronic Engineering,Computer Science Applications,Control and Systems Engineering

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A bi-annotated Malay-English code-switching (Manglish) dataset of X posts for biological gender identification and authorship attribution;Data in Brief;2024-02

2. Biological gender identification in Turkish news text using deep learning models;Multimedia Tools and Applications;2023-11-08

3. Topic Classification of Online News Articles Using Optimized Machine Learning Models;Computers;2023-01-09