Predicting Twitter User Demographics using Distant Supervision from Website Traffic Data-Reference-Cited by-同舟云学术

Predicting Twitter User Demographics using Distant Supervision from Website Traffic Data

Published:2016-02-19 Issue: Volume:55 Page:389-408
ISSN:1076-9757
Container-title:Journal of Artificial Intelligence Research
language:
Short-container-title:jair

Author:

Culotta Aron,Ravi Nirmal Kumar,Cutler Jennifer

Abstract

Understanding the demographics of users of online social networks has important applications for health, marketing, and public messaging. Whereas most prior approaches rely on a supervised learning approach, in which individual users are labeled with demographics for training, we instead create a distantly labeled dataset by collecting audience measurement data for 1,500 websites (e.g., 50% of visitors to gizmodo.com are estimated to have a bachelor's degree). We then fit a regression model to predict these demographics from information about the followers of each website on Twitter. Using patterns derived both from textual content and the social network of each user, our final model produces an average held-out correlation of .77 across seven different variables (age, gender, education, ethnicity, income, parental status, and political preference). We then apply this model to classify individual Twitter users by ethnicity, gender, and political preference, finding performance that is surprisingly competitive with a fully supervised approach.

Publisher

AI Access Foundation

Subject

Artificial Intelligence

Cited by 41 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Predicting hurricane evacuation behavior synthesizing data from travel surveys and social media;Transportation Research Part C: Emerging Technologies;2024-08

2. ExaAUAC: Arabic Twitter user age prediction corpus based on language and metadata features;Discover Artificial Intelligence;2024-07-08

3. Predicting the demographics of Twitter users with programmatic weak supervision;TOP;2024-02-08

4. Age-related bias and artificial intelligence: a scoping review;Humanities and Social Sciences Communications;2023-08-17

5. Neural age screening on question answering communities;Engineering Applications of Artificial Intelligence;2023-08