Abstract

In observational studies, high-dimensional multi-site data are subject to heterogeneous per

vasive hidden confounders across sites, leading to substantial bias in distributed estimation. To address

this challenge, we propose a deconfounded-debiased distributed estimation method, which integrates

majority voting for variable selection and the privacy-preserving aggregation of local deconfoundeddebiased estimators at the central site. This approach corrects biases from both heterogeneous per-

vasive hidden confounders and high-dimensional estimation, while also accommodating site-specific

heteroscedasticity in the random errors.

Theoretically, we prove that local deconfounded-debiased

estimators are asymptotically normal, and local individual hypothesis tests are asymptotically valid

with an asymptotic lower bound on their power. Furthermore, we establish variable selection consistency and asymptotic normality for the proposed distributed estimator. We also provide finite-sample

guarantees for variable selection, including an upper bound on the expected false positive rate and

a lower bound on the expected true positive rate. Simulation experiments and an application to a

protein dataset from the UK Biobank demonstrate our method’s superior finite-sample performance.

Key words and phrases: Distributed estimation, heterogeneous pervasive hidden confounders, majority voting, spectral transformation

Information

Preprint No.SS-2026-0052
Manuscript IDSS-2026-0052
Complete AuthorsZhaoyang Li, Chen Huang, Guoyou Qin, Zhongyi Zhu
Corresponding AuthorsGuoyou Qin
Emailsgyqin@fudan.edu.cn

References

  1. Bühlmann, P. and S. Van De Geer (2011). Statistics for high-dimensional data: methods, theory and applications. Springer Science & Business Media.
  2. Bycroft, C., C. Freeman, D. Petkova, G. Band, L. T. Elliott, K. Sharp, A. Motyer, D. Vukcevic, O. Delaneau,
  3. J. O’Connell, et al. (2018). The uk biobank resource with deep phenotyping and genomic data. Nature 562(7726), 203–209.
  4. Ćevid, D., P. Bühlmann, and N. Meinshausen (2020). Spectral deconfounding via perturbed sparse linear models. The Journal of Machine Learning Research 21(1), 9442–9482.
  5. Chen, X. and M. Xie (2014). A split-and-conquer approach for analysis of extraordinarily large data. Statistica Sinica 24(4), 1655–1684.
  6. Egashira, Y., H. Zhao, Y. Hua, R. F. Keep, and G. Xi (2015). White matter injury after subarachnoid hemorrhage. Stroke 46(10), 2909–2915.
  7. Fan, J., Y. Guo, and K. Wang (2023). Communication-efficient accurate statistical estimation. Journal of the American Statistical Association 118(542), 1000–1010.
  8. Fan, J., H. Liu, and W. Wang (2018). Large covariance estimation through elliptical factor models. Annals of statistics 46(4), 1383.
  9. Fan, J., Z. Lou, and M. Yu (2024). Are latent factor regression and sparse regression adequate? Journal of the American Statistical Association 119(546), 1076–1088.
  10. Festa, L. K., J. B. Grinspan, and K. L. Jordan-Sciutto (2024). White matter injury across neurodegenerative disease. Trends in Neurosciences 47(1), 47–57.
  11. Gao, Y., W. Liu, H. Wang, X. Wang, Y. Yan, and R. Zhang (2022). A review of distributed statistical inference. Statistical Theory and Related Fields 6(2), 89–99.
  12. Gold, D., J. Lederer, and J. Tao (2020). Inference for high-dimensional instrumental variables regression. Journal of Econometrics 217(1), 79–111.
  13. Gu, J. and S. X. Chen (2023). Distributed statistical inference under heterogeneity. Journal of Machine Learning Research 24(387), 1–57.
  14. Guo, Z., D. Ćevid, and P. Bühlmann (2022). Doubly debiased lasso: High-dimensional inference under hidden confounding. Annals of statistics 50(3), 1320.
  15. Hou, Z., W. Ma, and L. Wang (2023). Sparse and debiased lasso estimation and inference for high-dimensional composite quantile regression with distributed data. TEST 32(4), 1230–1250.
  16. Huang, H., W. Song, P. Wang, Y. Zhu, L. Zheng, C. Shen, H. Xu, and J. Qiu (2025). White matter hyperintensities: Cerebral small-vessel diseases and white matter microstructural impairments. iRADIOLOGY 3(1), 5–25.
  17. Javanmard, A. and A. Montanari (2014). Confidence intervals and hypothesis testing for high-dimensional regression. The Journal of Machine Learning Research 15(1), 2869–2909.
  18. Jickling, G. C., B. P. Ander, X. Zhan, B. Stamova, H. Hull, C. DeCarli, and F. R. Sharp (2022). Progression of cerebral white matter hyperintensities is related to leucocyte gene expression. Brain 145(9), 3179–3186.
  19. Jordan, M. I., J. D. Lee, and Y. Yang (2019). Communication-efficient distributed statistical inference. Journal of the American Statistical Association 114(526), 668–681.
  20. Li, Z., Y. Liu, K. Wei, Y. Yu, G. Qin, and Z. Zhu (2025). Deconfounded and debiased estimation for high-dimensional linear regression under hidden confounding with application to omics data. Bioinformatics 41(7), btaf400.
  21. Liu, W., X. Mao, and J. Tu (2025). Communication-efficient distributed sparse learning with oracle property and geometric convergence. JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION 120(552), 2606– 2618. Ogier du Terrail, J., Q. Klopfenstein, H. Li, I. Mayer, N. Loiseau, M. Hallal, M. Debouver, T. Camalon, T. Fouqueray,
  22. J. Arellano Castro, et al. (2025). Fedeca: federated external control arms for causal inference with time-to-event data in distributed settings. Nature Communications 16(1), 7496.
  23. Raudvere, U., L. Kolberg, I. Kuzmin, T. Arak, P. Adler, H. Peterson, and J. Vilo (2019). g: Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic acids research 47(W1), W191–W198.
  24. Shamir, O., N. Srebro, and T. Zhang (2014). Communication-efficient distributed optimization using an approximate newton-type method. In International conference on machine learning, pp. 1000–1008. PMLR.
  25. Sliz, E., J. Shin, S. Ahmad, D. M. Williams, S. Frenzel, F. Gauß, S. E. Harris, A.-K. Henning, M. V. Hernandez, Y.-H. Hu, B. Jiménez, M. Sargurupremraj, C. Sudre, R. Wang, K. Wittfeld, Q. Yang, J. M. Wardlaw, H. Völzke, M. W. Vernooij, J. M. Schott, M. Richards, P. Proitsi, M. Nauck, M. R. Lewis, L. Launer, N. Hosten, H. J.
  26. Grabe, M. Ghanbari, I. J. Deary, S. R. Cox, N. Chaturvedi, J. Barnes, J. I. Rotter, S. Debette, M. A. Ikram,
  27. M. Fornage, T. Paus, S. Seshadri, Z. Pausova, and for the NeuroCHARGE Working Group (2022). Circulating metabolome and white matter hyperintensities in women and men. Circulation 145(14), 1040–1052.
  28. Soldan, A., C. Pettigrew, Y. Zhu, M.-C. Wang, A. Moghekar, R. F. Gottesman, B. Singh, O. Martinez, E. Fletcher,
  29. C. DeCarli, et al. (2020). White matter hyperintensities and csf alzheimer disease biomarkers in preclinical alzheimer disease. Neurology 94(9), e950–e960.
  30. Sudlow, C., J. Gallacher, N. Allen, V. Beral, P. Burton, J. Danesh, P. Downey, P. Elliott, J. Green, M. Landray,
  31. et al. (2015). Uk biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS medicine 12(3), e1001779.
  32. Sun, T. and C.-H. Zhang (2012). Scaled sparse linear regression. Biometrika 99(4), 879–898.
  33. Sun, Y., L. Ma, and Y. Xia (2024). A decorrelating and debiasing approach to simultaneous inference for highdimensional confounded models. Journal of the American Statistical Association 119(548), 2857–2868.
  34. Tang, L., L. Zhou, and P. X.-K. Song (2020). Distributed simultaneous inference in generalized linear models via confidence distribution. Journal of Multivariate Analysis 176, 104567.
  35. Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society Series B: Statistical Methodology 58(1), 267–288.
  36. Van de Geer, S., P. Bühlmann, Y. Ritov, and R. Dezeure (2014). On asymptotically optimal confidence regions and tests for high-dimensional models. The Annals of Statistics 42(3), 1166–1202.
  37. Wang, J., Q. Zhao, T. Hastie, and A. B. Owen (2017). Confounder adjustment in multiple hypothesis testing. Annals of statistics 45(5), 1863.
  38. Wang, K. (2021). Unified distributed robust regression and variable selection framework for massive data. Expert Systems with Applications 186, 115701.
  39. Wartolowska, K. A. and A. J. Webb (2021). Blood pressure determinants of cerebral white matter hyperintensities and microstructural injury: Uk biobank cohort study. Hypertension 78(2), 532–539.
  40. Yang, S. and P. Ding (2020). Combining multiple observational data sources to estimate causal effects. Journal of the American Statistical Association 115(531), 1540–1554.
  41. Yu, G. and J. Bien (2019). Estimating the error variance in a high-dimensional linear model. Biometrika 106(3), 533–546.
  42. Yu, J., H. Wang, M. Ai, and H. Zhang (2022). Optimal distributed subsampling for maximum quasi-likelihood estimators with massive data. Journal of the American Statistical Association 117(537), 265–276.
  43. Yu, M., J. Li, and Y. Zhou (2026). Enhancements of communication-efficient distributed statistical inference and its privacy preservation. Journal of Econometrics 253, 106125.
  44. Zhang, C.-H. and S. S. Zhang (2014). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society: Series B: Statistical Methodology, 217–242.
  45. Zhao, H. and X. Shen (2025). Distributed algorithms for high-dimensional statistical inference and structure learning with heterogeneous data. Statistica Sinica 38(1). Zhaoyang Li

Acknowledgments

This work was supported by National Natural Science Foundation of China (No.82473724

to GQ).