Abstract

When confronted with features subject to outliers or asymmetry that

cause deviations from Gaussianity, classification remains a challenging task in

statistical learning.

We develop novel classification methods robust to Gaussian assumptions by adopting a flexible feature distribution, which is further

exploited to construct transformed features used in regularized logistic regression to achieve reliable binary classification and perform feature selection. We

investigate theoretical properties of the proposed classifier under realistic conditions that allow model misspecification, and establish consistency in separating

signal and noise features. Superior robustness and accuracy in classification of

the proposed methods compared to many existing methods are demonstrated in

simulation experiments, where we generate synthetic data representing a wide

range of distributional characteristics of features and data patterns. We showcase

the rich interpretable information beyond class labels that the proposed method

extracts from data in three real-life case studies.

Key words and phrases: heavy-tailed distribution, LASSO, likelihood ratio, logistic regression, mode

Information

Preprint No.SS-2026-0193
Manuscript IDSS-2026-0193
Complete AuthorsXiuchuan Liu, Xianzheng Huang
Corresponding AuthorsXiuchuan Liu
Emailsxiuchuan@email.sc.edu

References

  1. Arslan, O. (2012). Weighted LAD-LASSO method for robust parameter estimation and variable selection in regression. Computational Statistics & Data Analysis 56(6), 1952–1965.
  2. Azzalini, A. and A. Capitanio (2003). Distributions generated by perturbation of symmetry with emphasis on a multivariate skew t-distribution. Journal of the Royal Statistical Society: Series B 65(2), 367–389.
  3. Basu, A., A. Ghosh, M. Jaenada, and L. Pardo (2024). Robust adaptive LASSO in highdimensional logistic regression. Statistical Methods & Applications 33(5), 1217–1249.
  4. Breiman, L. (2001). Random forests. Machine learning 45(1), 5–32.
  5. Christen, P., D. J. Hand, and N. Kirielle (2023). A review of the f-measure: its history, properties, criticism, and alternatives. ACM Computing Surveys 56(3), 1–24.
  6. Duda, R., P. Hart, and D. Stork (2012). Pattern Classification. Wiley.
  7. Fan, J. and J. Lv (2008). Sure independence screening for ultrahigh dimensional feature space. Journal of the Royal Statistical Society Series B: Statistical Methodology 70(5), 849–911.
  8. Fern´andez, C. and M. F. Steel (1998). On Bayesian modeling of fat tails and skewness. Journal of the American Statistical Association 93(441), 359–371.
  9. Friedman, J. H., T. Hastie, and R. Tibshirani (2010). Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software 33, 1–22.
  10. Hampel, F. (1986). Robust Statistics: The Approach Based on Influence Functions. Probability and Statistics Series. Wiley.
  11. James, G., D. Witten, T. Hastie, R. Tibshirani, and J. Taylor (2023). An Introduction to Statistical Learning: with Applications in Python. Springer.
  12. Kahn, S. E., R. L. Hull, and K. M. Utzschneider (2006). Mechanisms linking obesity to insulin resistance and type 2 diabetes. Nature 444(7121), 840–846.
  13. Ke, G., Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems 30.
  14. Liu, Q., X. Huang, and R. Bai (2024). Bayesian modal regression based on mixture distributions. Computational Statistics & Data Analysis 199, 108012.
  15. Lundberg, S. M. and S.-I. Lee (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems 30.
  16. Lyon, R. J., B. W. Stappers, S. Cooper, J. M. Brooke, and J. D. Knowles (2016). Fifty years of pulsar candidate selection: from simple filters to a new principled real-time classification approach. Monthly Notices of the Royal Astronomical Society 459(1), 1104–1123.
  17. Park, M. Y. and T. Hastie (2007). L1-regularization path algorithm for generalized linear models. Journal of the Royal Statistical Society. Series B (Statistical Methodology) 69(4), 659–677.
  18. Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 1(5), 206–215.
  19. Sch¨olkopf, B. and A. J. Smola (2002). Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT press.
  20. Scrucca, L., C. Fraley, T. Murphy, and A. Raftery (2023). Model-Based Clustering, Classification, and Density Estimation Using mclust in R. Chapman & Hall/CRC The R Series. CRC Press.
  21. Singh, D., P. G. Febbo, K. Ross, D. G. Jackson, J. Manola, C. Ladd, P. Tamayo, A. A. Renshaw,
  22. A. V. D’Amico, J. P. Richie, et al. (2002). Gene expression correlates of clinical prostate cancer behavior. Cancer Cell 1(2), 203–209.
  23. Smith, J. W., J. E. Everhart, W. C. Dickson, W. C. Knowler, and R. S. Johannes (1988). Using the adap learning algorithm to forecast the onset of diabetes mellitus. In Proceedings of the Annual Symposium on Computer Application in Medical Care, pp. 261.
  24. Sun, Q., W.-X. Zhou, and J. Fan (2020). Adaptive Huber regression. Journal of the American Statistical Association 115(529), 254–265.
  25. Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society Series B: Statistical Methodology 58(1), 267–288.
  26. World Health Organization (2000). Obesity: preventing and managing the global epidemic: report of a WHO consultation.
  27. Xiong, W., W. K. H¨ardle, J. Wang, K. Yu, and M. Tian (2025). Mode-based classifier: A robust and flexible discriminant analysis for high-dimensional data. Statistica Sinica 35, 1391–1422.
  28. Zhao, P. and B. Yu (2006). On model selection consistency of Lasso. The Journal of Machine Learning Research 7, 2541–2563.

Supplementary Materials

include Appendices A–D for proofs referenced in

Section 3, Appendices E and F for implementation and simulation referenced

in Section 4, and Appendix G complementing discussions in Section 6.1.

Table 4: Average classification AUC and average feature selection F1 score across

300 Monte Carlo replicates for L1-RL and six variants of S-RoLLR in the six

simulation settings with continuous features, along with maximum AUC shortfall

and median runtime across the six settings

Setting

L1-RL Gaussian Student-t skew-normal skew-t TPSC mixture

AUC for classification (with the highest in each setting highlighted in boldface)


Supplementary materials are available for download.