Abstract: Reproducible learning of the underlying structure among large-scale network data is important in many contemporary applications. Despite the fast-growing literature on this subject, the practical issue of data heterogeneity has rarely been addressed. In this paper, we propose a new method called the multiple graphical knockoff filter to efficiently recover the underlying sparse connected structure of a general population from a high-dimensional heterogeneous dataset. We provide theoretical justification on the asymptotic false discovery rate control, and the theory for the power analysis is also established. To the best of our knowledge, this is the first formal theoretical result on the power for the graphical knockoffs procedure. Our new methodology and results are evidenced by numerical studies.
Key words and phrases: False discovery rate, heterogeneity, high-dimensionality, multiple graphical models, power.