Correcting a nonparametric two-sample graph hypothesis test for graphs with different numbers of vertices with applications to connectomics-Reference-Cited by-同舟云学术

Correcting a nonparametric two-sample graph hypothesis test for graphs with different numbers of vertices with applications to connectomics

Published:2024-01-03 Issue:1 Volume:9 Page:
ISSN:2364-8228
Container-title:Applied Network Science
language:en
Short-container-title:Appl Netw Sci

Author:

Alyakin Anton A.,Agterberg Joshua,Helm Hayden S.,Priebe Carey E.

Abstract

AbstractRandom graphs are statistical models that have many applications, ranging from neuroscience to social network analysis. Of particular interest in some applications is the problem of testing two random graphs for equality of generating distributions. Tang et al. (Bernoulli 23:1599–1630, 2017) propose a test for this setting. This test consists of embedding the graph into a low-dimensional space via the adjacency spectral embedding (ASE) and subsequently using a kernel two-sample test based on the maximum mean discrepancy. However, if the two graphs being compared have an unequal number of vertices, the test of Tang et al. (Bernoulli 23:1599–1630, 2017) may not be valid. We demonstrate the intuition behind this invalidity and propose a correction that makes any subsequent kernel- or distance-based test valid. Our method relies on sampling based on the asymptotic distribution for the ASE. We call these altered embeddings the corrected adjacency spectral embeddings (CASE). We also show that CASE remedies the exchangeability problem of the original test and demonstrate the validity and consistency of the test that uses CASE via a simulation study. Lastly, we apply our proposed test to the problem of determining equivalence of generating distributions in human connectomes extracted from diffusion magnetic resonance imaging at different scales.

Funder

Defense Advanced Research Programs Agency

Microsoft Research

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s41109-023-00607-x.pdf

Reference65 articles.

1. Agterberg J, Tang M, Priebe CE (2020) On two distinct sources of nonidentifiability in latent position random graph models. arXiv:2003.14250

2. Airoldi EM, Blei DM, Fienberg SE, Xing EP (2008) Mixed membership stochastic blockmodels. J Mach Learn Res 9:1981–2014

3. Arroyo J, Athreya A, Cape J, Chen G, Priebe CE, Vogelstein JT (2021) Inference for multiple heterogeneous networks with a common invariant subspace

4. Asta DM, Shalizi CR (2015) Geometric network comparisons. In: Proceedings of the thirty-first conference on uncertainty in artificial intelligence, UAI’15, Arlington, Virginia, United States, pp. 102–110. AUAI Press

5. Athreya A, Priebe CE, Tang M, Lyzinski V, Marchette DJ, Sussman DL (2016) A limit theorem for scaled eigenvectors of random dot product graphs. Sankhya A 78(1):1–18