Author:
Jamatia Anupam,Das Amitava,Gambäck Björn
Abstract
Abstract
This article addresses language identification at the word level in Indian social media corpora taken from Facebook, Twitter and WhatsApp posts that exhibit code-mixing between English-Hindi, English-Bengali, as well as a blend of both language pairs. Code-mixing is a fusion of multiple languages previously mainly associated with spoken language, but which social media users also deploy when communicating in ways that tend to be rather casual. The coarse nature of code-mixed social media text makes language identification challenging. Here, the performance of deep learning on this task is compared to feature-based learning, with two Recursive Neural Network techniques, Long Short Term Memory (LSTM) and bidirectional LSTM, being contrasted to a Conditional Random Fields (CRF) classifier. The results show the deep learners outscoring the CRF, with the bidirectional LSTM demonstrating the best language identification performance.
Subject
Artificial Intelligence,Information Systems,Software
Reference100 articles.
1. Conditional random fields: probabilistic models for segmenting and labeling sequence data,June 2001
2. The virtual speech community: social network and language variation on IRC;J. Comput. Mediat. Commun.,1999
3. Framewise phoneme classification with bidirectional LSTM and other neural network architectures,;Neural Netw.,2005
4. Code switch point detection in Arabic,2013
5. French in urban Lubumbashi Swahili: codeswitching, borrowing, or both?;J. Multiling. Multicult. Dev.,1992
Cited by
20 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献