Block models and personalized PageRank-Reference-Cited by-同舟云学术

Block models and personalized PageRank

Published:2016-12-20 Issue:1 Volume:114 Page:33-38
ISSN:0027-8424
Container-title:Proceedings of the National Academy of Sciences
language:en
Short-container-title:Proc Natl Acad Sci USA

Author:

Kloumann Isabel M.,Ugander Johan,Kleinberg Jon

Abstract

Methods for ranking the importance of nodes in a network have a rich history in machine learning and across domains that analyze structured data. Recent work has evaluated these methods through the “seed set expansion problem”: given a subsetSof nodes from a community of interest in an underlying graph, can we reliably identify the rest of the community? We start from the observation that the most widely used techniques for this problem, personalized PageRank and heat kernel methods, operate in the space of “landing probabilities” of a random walk rooted at the seed set, ranking nodes according to weighted sums of landing probabilities of different length walks. Both schemes, however, lack an a priori relationship to the seed set objective. In this work, we develop a principled framework for evaluating ranking methods by studying seed set expansion applied to the stochastic block model. We derive the optimal gradient for separating the landing probabilities of two classes in a stochastic block model and find, surprisingly, that under reasonable assumptions the gradient is asymptotically equivalent to personalized PageRank for a specific choice of the PageRank parameterαthat depends on the block model parameters. This connection provides a formal motivation for the success of personalized PageRank in seed set expansion and node ranking generally. We use this connection to propose more advanced techniques incorporating higher moments of landing probabilities; our advanced methods exhibit greatly improved performance, despite being simple linear classification rules, and are even competitive with belief propagation.

Funder

Simons Foundation

DOD | Army Research Office

Facebook

Google

David Morgenthaler

Publisher

Proceedings of the National Academy of Sciences

Subject

Multidisciplinary

Reference39 articles.

1. Page L Brin S Motwani R Winograd T (1998) The PageRank citation ranking: Bringing order to the web. Technical Report (Stanford InfoLab, Stanford, CA).

2. Authoritative sources in a hyperlinked environment

3. Gupta P (2013) WTF: The who to follow service at Twitter. Proceedings of the 22nd International Conference on World Wide Web (International World Wide Web Conference Committee, Geneva), pp 505–514.

4. PageRank beyond the web;Gleich;SIAM Rev,2015

5. Andersen R Lang KJ (2006) Communities from seed sets. Proceedings of the 15th International Conference on World Wide Web (American Association for Computing Machinery, New York), pp 223–232.

Cited by 53 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Bottom-up k-Vertex Connected Component Enumeration by Multiple Expansion;2024 IEEE 40th International Conference on Data Engineering (ICDE);2024-05-13

2. Estimating scope 3 greenhouse gas emissions through the shareholder network of publicly traded firms;Sustainability Science;2024-02-22

3. Anomaly Detection in Machining Centers Based on Graph Diffusion-Hierarchical Neighbor Aggregation Networks;Applied Sciences;2023-12-02

4. Starling: Introducing a mesoscopic scale with Confluence for Graph Clustering;PLOS ONE;2023-08-24

5. Modeling the Neurocognitive Dynamics of Language across the Lifespan;2023-07-04