Author:
Roberts Hal,Bhargava Rahul,Valiukas Linas,Jen Dennis,Malik Momin M.,Bishop Cindy Sherman,Ndulue Emily B.,Dave Aashka,Clark Justin,Etling Bruce,Faris Robert,Shah Anushka,Rubinovitz Jasmin,Hope Alexis,D'Ignazio Catherine,Bermejo Fernando,Benkler Yochai,Zuckerman Ethan
Abstract
We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.
Publisher
Association for the Advancement of Artificial Intelligence (AAAI)
Cited by
15 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献