Media Cloud: Massive Open Source Collection of Global News on the Open Web-Reference-Cited by-同舟云学术

Media Cloud: Massive Open Source Collection of Global News on the Open Web

Published:2021-05-22 Issue: Volume:15 Page:1034-1045
ISSN:2334-0770
Container-title:Proceedings of the International AAAI Conference on Web and Social Media
language:
Short-container-title:ICWSM

Author:

Roberts Hal,Bhargava Rahul,Valiukas Linas,Jen Dennis,Malik Momin M.,Bishop Cindy Sherman,Ndulue Emily B.,Dave Aashka,Clark Justin,Etling Bruce,Faris Robert,Shah Anushka,Rubinovitz Jasmin,Hope Alexis,D'Ignazio Catherine,Bermejo Fernando,Benkler Yochai,Zuckerman Ethan

Abstract

We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Cited by 15 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Lost in Recursion: Mining Rich Event Semantics in Knowledge Graphs;ACM Web Science Conference;2024-05-21

2. Epistemic language in news headlines shapes readers’ perceptions of objectivity;Proceedings of the National Academy of Sciences;2024-05-06

3. Comparing the Usage of Russian-and Ukrainian-Derived Search Terms to Evaluate the Impact of Misinformation, Disinformation, and Propaganda in the US;2024

4. Communication and democratic erosion: The rise of illiberal public spheres;European Journal of Communication;2023-12-25

5. Distinct information ecologies? Gender knowledge production in German digital legacy and counterpublic media;Feminist Media Studies;2023-11-09