Abstract
Accurate estimation of the extreme value index (EVI) is central to the analysis of tail
risks. Although substantial theoretical progress has been made, most existing estimators rely solely
on tail observations from the variable of interest, which are often too limited or unstable to yield
reliable inferences. We propose a novel framework, Individualized Fusion Learning (iFusion), to
enhance EVI estimation across multiple data sources. The core idea is to enhance estimation for a
target variable by strategically leveraging the information from the tails of related sources, thereby
improving efficiency while maintaining statistical validity. Under an independence assumption, we
construct an iFusion estimator based on the Hill estimator and establish its consistency and asymptotic normality. We extend the method to accommodate tail-dependent sources via an adjusted
fusion strategy. Simulation studies and an application to ozone concentration data demonstrate
the significantly improved performance of the iFusion method, particularly in variance reduction
and out-of-sample prediction, highlighting its immediate practical value for multi-source extreme
value analysis.
Key words and phrases: Asymptotic normality, Extreme value index, Individualized fusion learn- ing, Multi-source information borrowing, Variance reduction
Information
| Preprint No. | SS-2026-0109 |
|---|---|
| Manuscript ID | SS-2026-0109 |
| Complete Authors | Hongfang Sun, Yu Chen, Tao Xu, Zhengjun Zhang |
| Corresponding Authors | Zhengjun Zhang |
| Emails | zjz@stat.wisc.edu |
References
- Ahmed, H. and J. H. Einmahl (2019). Improved estimation of the extreme value index using related variables. Extremes 22(4), 553–569.
- Ahmed, H., J. H. Einmahl, and C. Zhou (2025). Extreme value statistics in semi-supervised models.
- Caeiro, F., M. I. Gomes, and D. Pestana (2005). Direct reduction of bias of the classical Hill estimator. REVSTAT-Statistical Journal 3(2), 113–136.
- Cai, C., R. Chen, and M.-g. Xie (2020). Individualized inference through fusion learning. Wiley Interdisciplinary Reviews: Computational Statistics 12(5), e1498. Chen, L., D. Li,
- and C. Zhou (2022). Distributed inference for the extreme value index. Biometrika 109(1), 257–264.
- Cui, Y. and M.-g. Xie (2023). Confidence distribution and distribution estimation for modern statistical inference. In Springer Handbook of Engineering Statistics, pp. 575–592. Springer.
- Daouia, A., S. A. Padoan, and G. Stupfler (2024). Optimal weighted pooling for inference about the tail index and extreme quantiles. Bernoulli 30(2), 1287–1312.
- de Haan, L. and A. Ferreira (2006). Extreme value theory: an introduction. Springer.
- de Haan, L. and C. Zhou (2021). Trends in extreme value indices. Journal of the American Statistical Association 116(535), 1265–1279.
- Drees, H. (2000). Weighted approximations of tail processes for β-mixing random variables. The Annals of Applied Probability 10(4), 1274–1301.
- Drees, H. and X. Huang (1998). Best attainable rates of convergence for estimators of the stable tail dependence function. Journal of Multivariate Analysis 64(1), 25–46.
- Duan, J., M. Pelger, and R. Xiong (2024). Target PCA: Transfer learning large dimensional panel data. Journal of Econometrics 244(2), 105521.
- Einmahl, J. H., L. de Haan, and C. Zhou (2016). Statistics of heteroscedastic extremes. Journal of the Royal Statistical Society Series B: Statistical Methodology 78(1), 31–51.
- Hill, B. M. (1975). A simple general approach to inference about the tail of a distribution. The Annals of Statistics 3(5), 1163–1174.
- Hsing, T. (1991). On tail index estimation using dependent data. The Annals of Statistics 19(3), 1547–1569.
- Jin, J., J. Yan, R. H. Aseltine, and K. Chen (2024). Transfer learning with large-scale quantile regression. Technometrics 66(3), 381–393.
- Kulik, R. and P. Soulier (2020). Heavy-tailed time series. Springer.
- Li, S., T. T. Cai, and H. Li (2022). Transfer learning for high-dimensional linear regression: Prediction, Methodology 84(1), 149–173.
- Li, S. and A. Luedtke (2023). Efficient estimation under data fusion. Biometrika 110(4), 1041–1054.
- Li, Y., L. Chen, D. Li, and H. Wang (2024). Estimating extreme value index by subsampling for massive datasets with heavy-tailed distributions. Statistics and Its Interface 17(4), 605–622.
- Shen, J., R. Y. Liu, and M.-g. Xie (2020). iFusion: Individualized fusion learning. Journal of the American Statistical Association 115(531), 1251–1267.
- Shi, X., Z. Pan, and W. Miao (2023). Data integration in causal inference. Wiley Interdisciplinary Reviews: Computational Statistics 15(1), e1581.
- Sun, B. and W. Miao (2022). On semiparametric instrumental variable estimation of average treatment effects through data fusion. Statistica Sinica 32, 569–590.
- Tian, Y. and Y. Feng (2023). Transfer learning under high-dimensional generalized linear models. Journal of the American Statistical Association 118(544), 2684–2697.
- Wang, H. and C.-L. Tsai (2009). Tail index regression. Journal of the American Statistical Association 104(487), 1233–1240.
- Wu, P., S. Luo, and Z. Geng (2025). On the comparative analysis of average treatment effects estimation via data combination. Journal of the American Statistical Association 120(552), 2250–2261.
- Xie, M.-g. and K. Singh (2013). Confidence distribution, the frequentist distribution estimator of a parameter: A review. International Statistical Review 81(1), 3–39.
- Zhang, Y. and Z. Zhu (2025). A data fusion method for quantile treatment effects. Statistica Sinica 35, 981–1002. Hongfang Sun, School of Mathematical Sciences, Ministry of Education Key Laboratory of NSLSCS, Nanjing Normal University
Acknowledgments
The authors acknowledge the Editor, the Associate Editor, and two anonymous reviewers for
their very helpful comments that led to a greatly improved version of this paper. The work
is supported by the National Natural Science Foundation of China (Nos. 12501657, 12371279,
12231017, 71991471, 72442027) and the Innovation Project (E5820801) of the Ministry of Education of China.
Supplementary Materials
The Supplementary Material contains the proofs of all theoretical results, auxiliary information,
as well as additional simulation studies and real data analyses.