Abstract

We propose a computationally efficient approach to performing maximum likelihood esti

mation of semiparametric models with big data containing potentially millions of observations. We

calculate an initial estimator for the vector of finite-dimensional parameters of primary interest using

a small fraction of the data and improve its efficiency by one-step estimation with semiparametric

efficient score functions on the remaining data. We estimate the efficient scores by the numerical differences of the profile log-likelihood function, eliminating the need for an analytic form of these scores,

which is often unavailable for semiparametric problems. We show that the resulting estimator is consistent and asymptotically normal with a covariance matrix that attains the semiparametric efficiency

bound. Because the majority of its computing time is spent in calculating the initial estimator using a

small subset of data, the proposed approach substantially reduces computational burden compared to

the standard method of performing maximum likelihood estimation on the entire dataset. We evaluate the performance of the proposed methods through extensive simulation studies on fitting the Cox

proportional hazards model to an interval-censored failure time and the partially linear model to a

continuous outcome. Finally, we apply the proposed methods to interval-censored data from the UK

Biobank.

Key words and phrases: Cox model, Efficient score, Interval-censored data, One-step estimation, Par- tially linear model, Profile likelihood

Information

Preprint No.SS-2026-0289
Manuscript IDSS-2026-0289
Complete AuthorsJianqiao Wang, Donglin Zeng, Danyu Lin
Corresponding AuthorsDanyu Lin
Emailslin@bios.unc.edu

References

  1. Bi, W., L. G. Fritsche, B. Mukherjee, S. Kim, and S. Lee (2020). A fast and accurate method for genome-wide timeto-event data analysis and its application to UK Biobank. The American Journal of Human Genetics 107(2), 222–233.
  2. Bickel, P. J., C. A. J. Klaassen, Y. Ritov, and J. A. Wellner (1993). Efficient and Adaptive Estimation for Semiparametric Models. Baltimore: Johns Hopkins University Press.
  3. Bycroft, C., C. Freeman, D. Petkova, G. Band, L. T. Elliott, K. Sharp, A. Motyer, D. Vukcevic, O. Delaneau,
  4. J. O’Connell, et al. (2018). The UK Biobank resource with deep phenotyping and genomic data. Nature 562, 203–209.
  5. Chen, X. (2007). Large sample sieve estimation of semi-nonparametric model. In J. J. Heckman and E. Leamer (Eds.), Handbook of Econometrics, Volume VI, pp. 5549–5632. New York: Elsevier Science B.V.
  6. Chen, Z., J. Chen, R. Collins, Y. Guo, R. Peto, F. Wu, and L. Li (2011). China kadoorie niobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. International Journal of Epidemiology 40(6), 1652–1666.
  7. Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal 21(1), C1–C68.
  8. Cox, D. R. (1972). Regression Models and Life-Tables. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 34 187 202
  9. Cox, D. R. (1975). Partial likelihood. Biometrika 62(2), 269–276.
  10. Dey, R., W. Zhou, T. Kiiskinen, A. Havulinna, A. Elliott, J. Karjalainen, M. Kurki, A. Qin, FinnGen, S. Lee, et al.
  11. (2022). Efficient and accurate frailty model approach for genome-wide survival association analysis in large-scale biobanks. Nature Communications 13(1), 5437.
  12. Gao, F., D. Zeng, D. Couper, and D. Lin (2019). Semiparametric regression analysis of multiple right-and intervalcensored events. Journal of the American Statistical Association 114(527), 1232–1240.
  13. Gaziano, J. M., J. Concato, M. Brophy, L. Fiore, S. Pyarajan, J. Breeling, S. Whitbourne, J. Deen, C. Shannon,
  14. D. Humphries, et al. (2016). Million veteran program: a mega-biobank to study genetic influences on health and disease. Journal of Clinical Epidemiology 70, 214–223.
  15. Geman, S. and C.-R. Hwang (1982). Nonparametric maximum likelihood estimation by the method of sieves. The Annals of Statistics 10, 401–414.
  16. Jordan, M. I., J. D. Lee, and Y. Yang (2019). Communication-efficient distributed statistical inference. Journal of the American Statistical Association 114(526), 668–681.
  17. Kalbfleisch, J. D. and R. L. Prentice (2002). The Statistical Analysis of Failure Time Data. New York: John Wiley & Sons.
  18. Lee, J. D., Q. Liu, Y. Sun, and J. E. Taylor (2017). Communication-efficient sparse regression. The American Journal of Human Genetics 18(5), 1–30.
  19. Lin, D.-Y. and D. Zeng (2010). On the relative efficiency of using summary statistics versus individual-level data in meta-analysis. Biometrika 97(2), 321–332.
  20. Murphy, S. A. and A. W. van der Vaart (2000). On profile likelihood. Journal of the American Statistical Association 95(450), 449–465.
  21. Neiswanger, W., C. Wang, and E. Xing (2014). Asymptotically exact, embarrassingly parallel MCMC. In Proceedings f th Thi ti th C f U t i t i A tifi i l I t lli A li t 623 632 AUAI P
  22. Ning, Y. and H. Liu (2017). A general theory of hypothesis tests and confidence regions for sparse high dimensional models. The Annals of Statistics 45(1), 158 – 195.
  23. Schifano, E. D., J. Wu, C. Wang, J. Yan, and M.-H. Chen (2016). Online updating of statistical inference in the big data setting. Technometrics 58(3), 393–403.
  24. Shen, X. and W. H. Wong (1994). Convergence rate of sieve estimates. The Annals of Statistics (2), 580–615.
  25. Thompson, D. J., D. Wells, S. Selzam, I. Peneva, R. Moore, K. Sharp, W. A. Tarran, E. J. Beard, F. Riveros-Mckay,
  26. C. Giner-Delgado, et al. (2022). Uk biobank release and systematic evaluation of optimised polygenic risk scores for 53 diseases and quantitative traits. MedRxiv, 2022–06.
  27. van der Vaart, A. W. (1998). Asymptotic Statistics. Cambridge: Cambridge University Press.
  28. van der Vaart, A. W. and J. A. Wellner (1996). Weak Convergence and Empirical Processes: With Applications to Statistics. New York: Springer.
  29. Wang, J., D. Zeng, and D. Y. Lin (2024). Fitting the Cox proportional hazards model to big data. Biometrics 80(1).
  30. Wang, W., S.-E. Lu, J. Q. Cheng, M. Xie, and J. B. Kostis (2022). Multivariate survival analysis in big data: a divide-and-combine approach. Biometrics 78(3), 852–866.
  31. Wang, Y., C. Hong, N. Palmer, Q. Di, J. Schwartz, I. Kohane, and T. Cai (2021, 09). A fast divide-and-conquer sparse Cox regression. Biostatistics 22(2), 381–401.
  32. Xu, Y., D. Zeng, and D. Y. Lin (2022, 11). Marginal proportional hazards models for multivariate interval-censored data. Biometrika 110(3), 815–830.
  33. Zeng, D. and D. Lin (2010). A general asymptotic theory for maximum likelihood estimation in semiparametric regression models with censored data. Statistica Sinica 20(2), 871.
  34. Zeng, D. and D. Y. Lin (2007). Maximum likelihood estimation in semiparametric regression models with censored data. Journal of the Royal Statistical Society Series B: Statistical Methodology 69(4), 507–564.
  35. Zeng, D., L. Mao, and D. Y. Lin (2016). Maximum likelihood estimation for semiparametric transformation models with interval-censored data. Biometrika 103, 253–271.
  36. Zhang, Y., J. C. Duchi, and M. J. Wainwright (2013). Communication-efficient algorithms for statistical optimization. Journal of Machine Learning Research 14(104), 3321–3363.
  37. Zhao, T., G. Cheng, and H. Liu (2016). A partially linear framework for massive heterogeneous data. The Annals of Statistics 44(4), 1400–1437.

Supplementary Materials

The Supplementary Material contains technical details of the proposed methods and additional simulation studies.


Supplementary materials are available for download.