Abstract
AbstractThe wealth of text data generated by social media has enabled new kinds of analysis of emotions with language models. These models are often trained on small and costly datasets of text annotations produced by readers who guess the emotions expressed by others in social media posts. This affects the quality of emotion identification methods due to training data size limitations and noise in the production of labels used in model development. We present LEIA, a model for emotion identification in text that has been trained on a dataset of more than 6 million posts with self-annotated emotion labels for happiness, affection, sadness, anger, and fear. LEIA is based on a word masking method that enhances the learning of emotion words during model pre-training. LEIA achieves macro-F1 values of approximately 73 on three in-domain test datasets, outperforming other supervised and unsupervised methods in a strong benchmark that shows that LEIA generalizes across posts, users, and time periods. We further perform an out-of-domain evaluation on five different datasets of social media and other sources, showing LEIA’s robust performance across media, data collection methods, and annotation schemes. Our results show that LEIA generalizes its classification of anger, happiness, and sadness beyond the domain it was trained on. LEIA can be applied in future research to provide better identification of emotions in text from the perspective of the writer.
Funder
Vienna Science and Technology Fund
European Research Council
H2020 European Research Council
Universität Konstanz
Publisher
Springer Science and Business Media LLC
Subject
Computational Mathematics,Computer Science Applications,Modeling and Simulation
Reference62 articles.
1. Pellert M, Schweighofer S, Garcia D (2021) Social media data in affective science. In: Handbook of computational social science, vol 1, pp 240–255. Routledge, London. https://doi.org/10.4324/9781003024583-18
2. De Choudhury M, Counts S, Gamon M (2012) Not all moods are created equal! Exploring human emotional states in social media. In: ICWSM, vol 6, pp 66–73
3. Golder SA, Macy MW (2011) Diurnal and seasonal mood vary with work, sleep, and daylength across diverse cultures. Science 333(6051):1878–1881
4. Garcia D, Rimé B (2019) Collective emotions and social resilience in the digital traces after a terrorist attack. Psychol Sci 30(4):617–628
5. Ferrara E, Yang Z (2015) Measuring emotional contagion in social media. PLoS ONE 10(11):0142390