Abstract

The estimation of causal effects based on the conditional average treatment effect (CATE) is usually vulnerable to outliers. However, to the best of our knowledge, outlier-resistant inference for the CATE has not been investigated in the literature. In this work, we propose an outlier-resistant estimation method for the CATE by incorporating M-estimation into the inverse propensity weighting (IPW) approach. The influence function and breakdown property are investigated to study the robustness of our method. In addition, we derive the asymptotic properties of the proposed estimator for inference purposes. The finite sample performance of the proposed estimator is evaluated via Monte Carlo experiments. The proposed method is compared with the IPW method and the augmented inverse probability weighting (AIPW) method, which do not account for outliers, as well as with DR-learner variants and machine-learning meta-learners, including the causal forest, X-learner, and R-learner. Finally, the proposed method is applied to the NHANES dataset to estimate the average effects of physical activity on albumin levels conditioned on age.

Key words and phrases: Conditional average treatment effect, Outlier-resistant estimators, Inverse probability weighting, M-estimator, Kernel regression

Information

Preprint No.SS-2026-0330
Manuscript IDSS-2026-0330
Complete AuthorsRan Mo, Honglang Wang
Corresponding AuthorsHonglang Wang
Emailshlwang@iu.edu

References

  1. Abrevaya, J., Y.-C. Hsu, and R. P. Lieli (2015). Estimating conditional average treatment effects. Journal of Business & Economic Statistics 33(4), 485–505.
  2. Angrist, J. D., G. W. Imbens, and D. B. Rubin (1996). Identification of causal effects using instrumental variables. J. Amer. Statist. Assoc. 91(434), 444–455.
  3. Athey, S. and G. Imbens (2016). Recursive partitioning for heterogeneous causal effects. Proc. Natl. Acad. Sci. USA 113(27), 7353–7360.
  4. Athey, S., J. Tibshirani, and S. Wager (2019). Generalized random forests. Ann. Statist. 47(2), 1148–1178.
  5. Beaton, A. E. and J. W. Tukey (1974). The fitting of power series, meaning polynomials, illustrated on bandspectroscopic data. Technometrics 16(2), 147–185.
  6. Chang, A., L. Van Horn, D. R. J. Jacobs, K. Liu, P. Muntner, and e. a. Newsome, Beth (2013). Lifestyle-related factors, obesity, and incident microalbuminuria: the cardia study. Am. J. Kidney Dis. 62(2), 267–275.
  7. Chipman, H. A., E. I. George, and R. E. McCulloch (2010). Bart: Bayesian additive regression trees. Ann. Appl. Stat 4(1) 266–298
  8. Cole, S. R. and M. A. Hern´an (2008). Constructing inverse probability weights for marginal structural models. American Journal of Epidemiology 168(6), 656–664.
  9. Comper, W. D. and T. M. Osicka (2005). Detection of urinary albumin. Advances in chronic kidney disease 12(2), 170–176.
  10. Cui, Y., M. R. Kosorok, E. Sverdrup, S. Wager, and R. Zhu (2023). Estimating heterogeneous treatment effects with right-censored data via causal survival forests. J. R. Stat. Soc. Ser. B. Stat. Methodol. 85(2), 179–211.
  11. Donoho, D. L. and P. J. Huber (1983). The notion of breakdown point. A festschrift for Erich L. Lehmann 157184.
  12. Dorie, V., J. Hill, U. Shalit, M. Scott, and D. Cervone (2019). Automated versus do-it-yourself methods for causal inference: Lessons learned from a data analysis competition. Statist. Sci. 34(1), 43–68.
  13. Fan, Q., Y.-C. Hsu, R. P. Lieli, and Y. Zhang (2022). Estimation of conditional average treatment effects with high-dimensional data. J. Bus. Econom. Statist. 40(1), 313–327.
  14. Firpo, S. (2007). Efficient semiparametric estimation of quantile treatment effects. Econometrica 75(1), 259–276.
  15. Foster, J. C., J. M. Taylor, and S. J. Ruberg (2011). Subgroup identification from randomized clinical trial data. Statistics in medicine 30(24), 2867–2880.
  16. Giloni, A. and J. S. Simonoff (2005). The conditional breakdown properties of least absolute value local polynomial estimators. Journal of Nonparametric Statistics 17(1), 15–30.
  17. Glynn, A. N. and K. M. Quinn (2010). An introduction to the augmented inverse propensity weighted estimator. Political analysis 18(1), 36–56.
  18. Hambrecht, R., E. Fiehn, C. Weigl, S. Gielen, C. Hamann, R. Kaiser, J. Yu, V. Adams, J. Niebauer, and G. Schuler
  19. (1998). Regular physical exercise corrects endothelial dysfunction and improves exercise capacity in patients with chronic heart failure. Circulation 98(24), 2709–2715.
  20. Hambrecht, R., A. Wolf, S. Gielen, A. Linke, J. Hofer, S. Erbs, N. Schoene, and G. Schuler (2000). Effect of exercise on coronary endothelial function in patients with coronary artery disease. New England Journal of Medicine 342(7), 454–460.
  21. Hampel, F. R. (1974). The influence curve and its role in robust estimation. Journal of the american statistical association 69(346), 383–393.
  22. Harada, K. and H. Fujisawa (2021). Outlier-resistant estimators for average treatment effect in causal inference. arXiv preprint arXiv:2106.13946.
  23. H¨ardle, W. (1984). Robust regression function estimation. Journal of Multivariate Analysis 14(2), 169–180.
  24. Hill, J. L. (2011). Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics 20(1), 217–240.
  25. Hirano, K., G. W. Imbens, and G. Ridder (2003). Efficient estimation of average treatment effects using the estimated propensity score. Econometrica 71(4), 1161–1189.
  26. Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association 81(396), 945–960.
  27. Holland, P. W. and R. E. Welsch (1977). Robust regression using iteratively reweighted least-squares. Communications in Statistics-theory and Methods 6(9), 813–827.
  28. Huber, P. J. (1964). Robust estimation of a location parameter. Ann. Math. Statist. 35(4), 73–101.
  29. Huber, P. J. (1973). Robust regression: asymptotics, conjectures and monte carlo. The annals of statistics, 799–821.
  30. Kennedy, E. H. (2023). Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics 17(2), 3008–3049.
  31. K¨unzel, S. R., J. S. Sekhon, P. J. Bickel, and B. Yu (2019). Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences 116(10), 4156–4165.
  32. Kuo, H.-Y., Y.-H. Huang, S.-W. Wu, F.-H. Chang, Y.-W. Tsuei, H.-C. Fan, and et al. (2022). The effects of exercise habit on albuminuria and metabolic indices in patients with type 2 diabetes mellitus: A cross-sectional study. Medicina 58(5), 577.
  33. Kurz, C. F. (2022). Augmented inverse probability weighting and the double robustness property. Medical Decision Making 42(2), 156–167.
  34. Lang, J., R. Katz, J. H. Ix, O. M. Gutierrez, C. A. Peralta, C. R. Parikh, S. Satterfield, S. Petrovic, P. Devarajan,
  35. M. Bennett, et al. (2018). Association of serum albumin levels with kidney function decline and incident chronic kidney disease in elders. Nephrology Dialysis Transplantation 33(6), 986–992.
  36. Lee, S., R. Okui, and Y.-J. Whang (2017). Doubly robust uniform confidence band for the conditional average treatment effect function. Journal of Applied Econometrics 32(7), 1207–1225.
  37. Lee, S. and Y.-J. Whang (2009). Nonparametric tests of conditional treatment effects. (No. 1740).
  38. Leon, A. S., J. Connett, D. R. Jacobs, and R. Rauramaa (1987). Leisure-time physical activity levels and risk of coronary heart disease and death: the multiple risk factor intervention trial. Jama 258(17), 2388–2395.
  39. Li, G. and J. Zhang (1998). Breakdown properties of location M-estimators. The Annals of Statistics 26(3), 1170– 1189.
  40. Luedtke, A. R. and M. J. van der Laan (2016). Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy. Annals of Statistics 44(2), 713–742.
  41. Luedtke, A. R. and M. J. van der Laan (2017). Evaluating the impact of treating the optimal subgroup. Statistical Methods in Medical Research 26(4), 1630–1640.
  42. Miao, W., Z. Geng, and E. J. Tchetgen Tchetgen (2018). Identifying causal effects with proxy variables of an unmeasured confounder Biometrika 105(4) 987 993
  43. Miller, W. G., D. E. Bruns, G. L. Hortin, S. Sandberg, K. M. Aakre, M. J. McQueen, Y. Itoh, J. C. Lieske, D. W.
  44. Seccombe, G. Jones, et al. (2009). Current issues in measurement and reporting of urinary albumin excretion. Clinical chemistry 55(1), 24–38.
  45. Neyman, J. (1923). Sur les applications de la th´eorie des probabilit´es aux experiences agricoles: Essai des principes. Roczniki Nauk Rolniczych 10(1), 1–51.
  46. Nie, X. and S. Wager (2021). Quasi-oracle estimation of heterogeneous treatment effects. Biometrika 108(2), 299–319.
  47. Powers, S., J. Qian, K. Jung, A. Schuler, N. H. Shah, T. Hastie, and R. Tibshirani (2018). Some methods for heterogeneous treatment effect estimation in high dimensions. Statistics in medicine 37(11), 1767–1787.
  48. Robins, J. M. and D. M. Finkelstein (2000). Correcting for noncompliance and dependent censoring in an AIDS clinical trial with inverse probability of censoring weighted (IPCW) log-rank tests. Biometrics 56(3), 779–788.
  49. Robins, J. M. and A. Rotnitzky (1995). Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Association 90(429), 122–129.
  50. Robins, J. M., A. Rotnitzky, and L. P. Zhao (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association 89(427), 846–866.
  51. Robins, J. M., M. Sued, Q. Lei-Gomez, and A. Rotnitzky (2007). Comment: Performance of double-robust estimators when “inverse probability” weights are highly variable. Statistical Science 22(4), 544–559.
  52. Robinson, E. S., N. D. Fisher, J. P. Forman, and G. C. Curhan (2010). Physical activity and albuminuria. American journal of epidemiology 171(5), 515–521.
  53. Rosenbaum, P. R. and D. B. Rubin (1983). The central role of the propensity score in observational studies for causal effects. Biometrika 70(1), 41–55.
  54. Rousseeuw, P. J., F. R. Hampel, E. M. Ronchetti, and W. A. Stahel (1986). Robust statistics: the approach based on influence functions
  55. Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology 66(5), 688.
  56. Ruppert, D., S. J. Sheather, and M. P. Wand (1995). An effective bandwidth selector for local least squares regression. Journal of the American Statistical Association 90(432), 1257–1270.
  57. Solbu, M. D., J. Kronborg, B. O. Eriksen, T. G. Jenssen, and I. Toft (2008). Cardiovascular risk-factors predict progression of urinary albumin-excretion in a general, non-diabetic population: a gender-specific follow-up study. Atherosclerosis 201(2), 398–406.
  58. Sverdrup, E. and Y. Cui (2023). Proximal causal learning of conditional average treatment effects. In Proceedings of the 40th International Conference on Machine Learning, Volume 202 of Proceedings of Machine Learning Research, pp. 33285–33298.
  59. van der Laan, M. J. (2006). Statistical inference for variable importance. International Journal of Biostatistics 2(1).
  60. VanderWeele, T. J., A. R. Luedtke, M. J. van der Laan, and R. C. Kessler (2019). Selecting optimal subgroups for treatment using many covariates. Epidemiology 30(3), 334–341.
  61. Yao, L., S. Li, Y. Li, M. Huai, J. Gao, and A. Zhang (2018). Representation learning for treatment effect estimation from observational data. Advances in neural information processing systems 31.
  62. Zhang, Z., Z. Chen, J. F. Troendle, and J. Zhang (2012). Causal inference on quantiles with an obstetric application. Biometrics 68(3), 697–706.

Acknowledgments

This research was supported in part by the National Science Foundation under Grant DMS-2212928 and by Lilly Endowment, Inc., through its support for the Indiana University Pervasive Technology Institute.

The study utilizes data from the National Health and Nutrition Examination Survey (NHANES), conducted by the National Center for Health Statistics, part of the Centers for Disease Control and Prevention (CDC).

Supplementary Materials

Technical proofs, detailed simulation results, and the full DR-learner and machine-learning meta-learner comparison studies summarised in Section 5.3, together with their implementation details, are provided in the supplementary material.


Supplementary materials are available for download.