Abstract
Feature selection and importance estimation in a model-agnostic setting is an ongoing chal
lenge of significant interest. Wrapper methods are commonly used because they are typically modelagnostic. In this paper, we develop a general comparison framework for model-agnostic feature selection
methods based on relative efficiency, using relative variability σ/µ to account for different statistics
having different means. In particular we focus on state-of-the-art feature selection methods, the Generalized Covariance Measure (GCM) and Leave-One-Covariate-Out (LOCO) estimation. In particular,
we present a theoretical comparison under three model settings: linear models, non-linear additive
models, and single index models that mimic a single-layer neural network. We complement this with
simulations and real data examples for the above models and mis-specified models. Our theoretical
results, along with empirical findings, demonstrate that GCM-related methods generally out-perform
LOCO under suitable regularity conditions defined by a suitably defined correlation quantity which
quantifies the asymptotic relative efficiency of these approaches. Our simulations and real data analysis
include widely used machine learning methods such as neural networks and gradient boosting trees.
Key words and phrases: conditional independence; feature selection; generalized covariance measure; leave-one-covariate-out; model-agnostic methods; variable importance
Information
| Preprint No. | SS-2025-0337 |
|---|---|
| Manuscript ID | SS-2025-0337 |
| Complete Authors | Chenghui Zheng, Garvesh Raskutti |
| Corresponding Authors | Chenghui Zheng |
| Emails | chenghui.zheng@wisc.edu |
References
- Altmann, A., L. Tolo¸si, O. Sander, and T. Lengauer (2010). Permutation importance: a corrected feature importance measure. Bioinformatics 26(10), 1340–1347.
- Balabdaoui, F., C. Durot, and H. Jankowski (2019). Least squares estimation in the monotone single index model. Bernoulli 25(4B).
- Barber, R. F., E. Candes, L. Janson, E. Patterson, and M. Sesia (2022). knockoff: The Knockoff Filter for Controlled Variable Selection.
- Barber, R. F. and E. J. Cand`es (2015). Controlling the False Discovery Rate Via Knockoffs. The Annals of Statistics 43(5), 2055–2085.
- Breiman, L. (2001). Random Forests. Machine Learning 45(1), 5–32.
- Cand`es, E., Y. Fan, L. Janson, and J. Lv (2018). Panning for Gold: ‘Model-X’ Knockoffs for High Dimensional Controlled Variable Selection. Journal of the Royal Statistical Society Series B: Statistical Methodology 80(3), 551–577.
- Chang, C.-H., L. Rampasek, and A. Goldenberg (2018). Dropout Feature Ranking for Deep Learning Models.
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2 ed.). Lawrence Erlbaum.
- Cortizo, J. C. and I. Giraldez (2006). Multi Criteria Wrapper Improvements to Naive Bayes Learning. In Intelligent
- Duda, R. O., P. Hart, and D. Stork (2000). Pattern classification/Duda RO, Hart PE, Stork DG–.
- Fisher, A., C. Rudin, and F. Dominici (2019). All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research 20(177), 1–81.
- Fisher, R. A. (1992). Statistical Methods for Research Workers. In S. Kotz and N. L. Johnson (Eds.), Breakthroughs in Statistics: Methodology and Distribution, pp. 66–70. Springer.
- Fleuret, F. (2004). Fast binary feature selection with conditional mutual information. Journal of Machine learning research 5(9).
- Friedman, J., T. Hastie, R. Tibshirani, B. Narasimhan, K. Tay, N. Simon, J. Qian, and J. Yang (2023). glmnet: Lasso and Elastic-Net Regularized Generalized Linear Models.
- Gao, Y., A. Stevens, G. Raskutti, and R. Willett (2022). Lazy Estimation of Variable Importance for Large Neural Networks. International Conference on Machine Learning, 7122–7143.
- Guyon, I., J. Weston, S. Barnhill, and V. Vapnik (2002). Gene Selection for Cancer Classification using Support Vector Machines. Machine Learning 46(1), 389–422.
- Hoque, N., D. K. Bhattacharyya, and J. K. Kalita (2014). MIFS-ND: A mutual information-based feature selection method. Expert Systems with Applications 41(14), 6371–6385.
- Inside Airbnb (2026). Inside airbnb: Get the data. https://insideairbnb.com/get-the-data/. Quebec City,
- Quebec, Canada listings dataset, June 2026.
- Kittler, J. (1978). Feature set search algorithms. Pattern recognition and signal processing.
- Lehmann, E. L. and J. P. Romano (2005). Testing Statistical Hypotheses (3 ed.). Springer.
- Lei, J., M. G’Sell, A. Rinaldo, R. J. Tibshirani, and L. Wasserman (2018). Distribution-Free Predictive Inference for Regression. Journal of the American Statistical Association 113(523), 1094–1111.
- Ma, S. and J. Huang (2008). Penalized feature selection and classification in bioinformatics. Briefings in Bioinformatics 9(5), 392–403.
- Meinshausen, N. and P. B¨uhlmann (2010). Stability selection. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 72(4), 417–473.
- Mentch, L. and G. Hooker (2016). Quantifying uncertainty in random forests via confidence intervals and hypothesis tests. Journal of Machine Learning Research 17(26), 1–41.
- Minai, A. A. and R. D. Williams (1993). On the derivatives of the sigmoid. Neural Networks 6(6), 845–853.
- Pearson, K. and F. Galton (1997). VII. Note on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London 58(347-352), 240–242.
- Ridgeway, G., D. Edwards, B. Kriegler, S. Schroedl, H. Southworth, B. Greenwell, B. Boehmke, J. Cunningham, and
- GBM Developers (2024). gbm: Generalized Boosted Regression Models.
- Rinaldo, A., L. Wasserman, and M. G’Sell (2019). Bootstrapping and sample splitting for high-dimensional, assumption-lean inference. The Annals of Statistics 47(6), 3438–3469.
- Sandri, M. and P. Zuccolotto (2006). Variable Selection Using Random Forests. In S. Zani, A. Cerioli, M. Riani, and M. Vichi (Eds.), Data Analysis, Classification and the Forward Search, pp. 263–270. Springer.
- Shah, R. D. and J. Peters (2020). The Hardness of Conditional Independence Testing and the Generalised Covariance Measure. The Annals of Statistics 48(3).
- Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology 15(1), 72–101.
- Strobl, C., A.-L. Boulesteix, T. Kneib, T. Augustin, and A. Zeileis (2008). Conditional variable importance for random forests. BMC Bioinformatics 9(1), 307.
- Sz´ekely, G. J., M. L. Rizzo, and N. K. Bakirov (2007). Measuring and testing dependence by correlation of distances.
- Tansey, W., V. Veitch, H. Zhang, R. Rabadan, and D. M. Blei (2022). The Holdout Randomization Test for Feature Selection in Black Box Models. Journal of Computational and Graphical Statistics.
- van der Vaart, A. W. (2000). Asymptotic Statistics. Cambridge University Press.
- Vangel, M. G. (1996). Confidence intervals for a normal coefficient of variation. The American Statistician 50(1), 21–26.
- Williamson, B. D., P. B. Gilbert, N. R. Simon, and M. Carone (2022). A General Framework for Inference on Algorithm-Agnostic Variable Importance. Journal of the American Statistical Association 118(543), 1645–1658.
- Witten, I. H., E. Frank, M. A. Hall, C. J. Pal, and M. Data (2005). Practical machine learning tools and techniques, Volume 2. Elsevier Amsterdam, The Netherlands. Issue: 4.
- Yu, L. and H. Liu (2003). Feature selection for high-dimensional data: A fast correlation-based filter solution. In Proceedings of the 20th international conference on machine learning (ICML-03), pp. 856–863.
- Zaman, A. (2026). Social media addiction vs productivity dataset.
- Zhang, K., J. Peters, D. Janzing, and B. Schoelkopf (2012). Kernel-based Conditional Independence Test and Application in Causal Discovery.
- Zou, H. and T. Hastie (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society Series B: Statistical Methodology 67(2), 301–320.
Supplementary Materials
The online Supplementary Material contains an extra efficiency comparison example for
additive model (Section S1). Mean and variance for test statistics under different models
(Section S2), proofs of theorems (Section S3), and additional simulation and data analysis
results (Section S4).