<?xml version='1.0' encoding='UTF-8' ?><rss version='2.0'><channel><title>Statistica Sinica: Volume 36, Number 3, July 2026</title><description>This is an example of an RSS feed</description><link>https://www3.stat.sinica.edu.tw/statistica/</link><lastBuildDate>Tue, 7 July 2026 00:01:00 +0000 </lastBuildDate><pubDate>Tue, 7 July 2026 00:01:00 +0000 </pubDate><ttl>1800</ttl>
<item>
<link>/statistica/J36N3/J36N301/J36N301.html</link>
<title> MODEL AVERAGING ESTIMATION FOR PARTIALLY LINEAR FUNCTIONAL SCORE MODELS </title>
<author>Shishi Liu, Chunming Zhang, Hao Zhang, Rou Zhong and Jingxiao Zhang </author>
<page>1043-1067. DOI:10.5705/ss.202023.0067</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; The scalar-on-function regression is quite useful for modelling mixed-data in the context of scalar and functional variables. Under this class of regression, the paper aims at proposing a compelling alternative to model selection methods to address model selection uncertainty. The considered models characterize a scalar response using parametric effect of the scalar predictors and nonparametric effect of a functional predictor, and a model averaging estimation is developed based on Mallows-type criterion to assign weights for averaging. Further, the asymptotic optimality of the resulting estimator, in terms of achieving the smallest possible squared error loss, is established. Besides, simulation studies demonstrate its superiority to or comparability with some information criterion score-based model selection and averaging estimators. The proposed procedure is also applied to a mid-infrared spectra dataset for illustration. &lt;p&gt;Key words and phrases: Functional data, mallows-type criterion, model average.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N302/j36N302.html</link>
<title> PRINCIPAL SUB-MANIFOLDS </title>
<author>Zhigang Yao, Benjamin Eltzner and Tung Pham </author>
<page>1069-1089. DOI:10.5705/ss.202021.0163</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; We propose a novel method of finding principal components in multivariate data sets that lie on an embedded nonlinear Riemannian manifold within a higher-dimensional space. Our aim is to extend the geometric interpretation of PCA, while being able to capture non-geodesic modes of variation in the data. We introduce the concept of a principal sub-manifold, a manifold passing through a reference point, and at any point on the manifold extending in the direction of highest variation in the space spanned by the eigenvectors of the local tangent space PCA. Compared to recent work for the case where the sub-manifold is of dimension one (Panaretos et al., 2014)-essentially a curve lying on the manifold attempting to capture one-dimensional variation-the current setting is much more general. The principal sub-manifold is therefore an extension of the principal flow, accommodating to capture higher dimensional variation in the data. We show the principal sub-manifold yields the ball spanned by the usual principal components in Euclidean space. By means of examples, we illustrate how to find, use and interpret a principal sub-manifold and we present an application in shape analysis. &lt;p&gt;Key words and phrases: Dimension reduction, manifold, principal component analysis, shape analysis, tangent space.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N303/j36N303.html</link>
<title> TESTING FOR THE EQUALITY OF DISTRIBUTIONS IN HIGH DIMENSION </title>
<author>Xu Li, Gongming Shi and Baoxue Zhang </author>
<page>1091-1111. DOI:10.5705/ss.202023.0299</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; In this paper, we propose a new homogeneous test for two highd-imensional random vectors. Our test is built on a new measure, the so-called characteristic distance, which can completely characterize the homogeneity of two distributions. The newly proposed metric has some desirable properties, for example, it possesses a clear and intuitive probabilistic interpretation, and can be used to address the high-dimensional distance inference. Theoretically, the limiting behaviors under the conventional fixed dimension and high-dimensional distance inference are thoroughly investigated. Simulation studies and real data analysis are presented to illustrate the finite-sample performance of the proposed test statistic. &lt;p&gt;Key words and phrases: Characteristic distance, high dimensionality, permutation procedure, test of homogeneity, U-statistic.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N304/j36N304.html</link>
<title> NEARLY OPTIMAL TWO-STEP POISSON SAMPLING AND EMPIRICAL LIKELIHOOD WEIGHTING ESTIMATION FOR M-ESTIMATION WITH BIG DATA </title>
<author>Yan Fan, Yang Liu, Yukun Liu and Jing Qin </author>
<page>1113-1132. DOI:10.5705/ss.202023.0274</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Subsampling techniques can effectively reduce the computational costs of processing big data. Practical subsampling plans typically involve initial uniform sampling and refined sampling. Subsample-based big data inferences are generally built on the inverse probability weighting (IPW), which may be unstable and cannot incorporate auxiliary information. In this paper, we consider a two-step Poisson sampling, which combines an initial uniform sampling with a second Poisson sampling. Under this sampling plan, we propose an empirical likelihood weighting (ELW) estimation approach to an M-estimation parameter, and then construct a nearly optimal two-step Poisson sampling plan based on the ELW method to improve estimation efficiency of IPW-based optimal subsamplings. Further, we derive methods for determining the smallest sample sizes with which the proposed sampling-and-estimation method produces estimators of guaranteed precision. Our ELW method overcomes the instability of IPW by circumventing the use of inverse probabilities, and utilizes auxiliary information including the size and certain sample moments of big data. We show that the proposed ELW method produces more efficient estimators than IPW, leading to more efficient optimal sampling plans and more economical sample sizes for a prespecified estimation precision. These advantages are confirmed through real data based simulations. &lt;p&gt;Key words and phrases: Big data, empirical likelihood, two-step Poisson sampling.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N305/j36N305.html</link>
<title> INTRINSIC MINIMUM AVERAGE VARIANCE ESTIMATION FOR DIMENSION REDUCTION WITH SYMMETRIC POSITIVE-DEFINITE MATRICES AND BEYOND </title>
<author>Baiyu Chen, Shuang Dai and Zhou Yu </author>
<page>1133-1153. DOI:10.5705/ss.202023.0268</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; In this paper, we estimate the central mean subspace in a dimension reduction problem where the response is a symmetric positive-definite matrix. We propose the intrinsic minimum average variance estimation and the intrinsic outer product of gradient method which fully exploit the geometric structure of the Riemannian manifold where the response resides. We present algorithms for our newly developed methods under the log-Euclidean metric and the log-Cholesky metric. The two metrics endow the manifold with a commutative Lie group structure that transforms our manifold model into a Euclidean one and helps us derive the consistency and asymptotic normality of estimators. Our methods are then naturally extended to the case allowing p = pn to diverge and the case of general Riemannian manifolds. Several simulation studies and an application to the New York taxi network data showcase the superiority of our proposals. &lt;p&gt;Key words and phrases: Central mean subspace, log-Cholesky metric, log-Euclidean metric, minimum average variance estimation, outer product of gradient, symmetric positive-definite matrix.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N306/j36N306.html</link>
<title> TWO KERNEL-BASED FEATURE SCREENING PROCEDURES FOR HIGH-DIMENSIONAL RESPONSE DATA </title>
<author>Yuke Shi, Na Li, Qizhai Li, Dongdong Pan and Jinjuan Wang </author>
<page>1155-1174. DOI:10.5705/ss.202023.0290</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; We consider feature screening for high-dimensional response data without and with the existence of confounding factors. First, we introduce kernel covariance and kernel correlation for high-dimensional associaiton analysis, and further propose partial kernel covariance and partial kernel correlation that can handle situations with confounding factors. Then, based on the kernel correlation and partial kernel correlation, we propose two feature screening procedures. Both screening procedures possess sure screening property and ranking consistency property, and are complementary to each other by respectively dealing with situations without and with the existence of confounding factors. The proposed procedures make no assumptions on model, and are suitable for high-dimensional response variable and non-Euclidean data. Extensive simulation results and a real data analysis demonstrate the satisfying performances and advantages of the proposed procedures over existing methods. &lt;p&gt;Key words and phrases: Confounding factors, feature screening, high-dimensional response variable, kernel correlation, partial kernel correlation.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N307/j36N307.html</link>
<title> BAYESIAN STATISTICS BY ARITHMETIC OPERATIONS OF CONJUGATE DISTRIBUTIONS </title>
<author>Hang Qian </author>
<page>1175-1192. DOI:10.5705/ss.202024.0052</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Conjugate distributions provide an entry point to Bayesian analysis. By defining summation, subtraction, and multiplication operators for conjugate distributions, we study Bayesian statistics by arithmetic operations. A striking feature is that the non-informative prior fulfills the central role of zero in mathematics. The summation operator connects Bayesian and frequentist estimators by a simple equation, which also provides an efficient method for evaluating the marginal likelihood. The subtraction operator facilitates cross-validation, rolling-window estimation, and regression under multicollinearity. The multiplication operator simplifies the weighted regression with a discount factor. Arithmetic operations conceptualize pseudo data in the conjugate prior, sufficient statistics that determine the likelihood, and the posterior that balances the prior and data. &lt;p&gt;Key words and phrases: Conjugacy, exponential family, linear regression, statistics education.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N308/j36N308.html</link>
<title> IDENTIFIABILITY AND ESTIMATION OF CAUSAL EFFECTS WITH NON-GAUSSIANITY AND AUXILIARY COVARIATES </title>
<author>Kang Shuai, Shanshan Luo, Yue Zhang, Feng Xie and Yangbo He </author>
<page>1193-1212. DOI:10.5705/ss.202023.0315</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Assessing causal effects in the presence of unmeasured confounding is challenging. Although auxiliary variables, such as instrumental variables, are commonly used to identify causal effects, they are often unavailable in practice due to stringent and untestable conditions. To address this issue, previous researches have utilized linear structural equation models to show that the causal effect is identifiable when noise variables of the treatment and outcome are both non-Gaussian. In this paper, we investigate the problem of identifying the causal effect using the auxiliary covariate and non-Gaussianity from the treatment. Our key idea is to characterize the impact of unmeasured confounders using an observed covariate, assuming they are all Gaussian. We demonstrate that the causal effect can be identified using a measured covariate, and then extend the identification results to the multi-treatment setting. We further develop a simple estimation procedure for estimating causal effects and derive a &amp;radic;n-consistent estimator. Finally, we evaluate the performance of our estimator through simulation studies and apply our method to investigate the effect of the trade on income. &lt;p&gt;Key words and phrases: Auxiliary variable, causal effects, identification, multiple treatments, non-Gaussianity.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N309/j36N309.html</link>
<title> LOCALLY OPTIMAL DESIGNS FOR ESTIMATING ONE OR MORE FUNCTIONS OF SHARED PARAMETERS BETWEEN TWO GROUPS IN BIOMEDICAL STUDIES </title>
<author>Xin Liu, Rong-Xian Yue and Weng Kee Wong </author>
<page>1213-1234. DOI:10.5705/ss.202023.0284</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Models with shared parameters arise quite naturally in the biological sciences and we use optimal design theory to construct c-optimal approximate designs for estimating one or more functions of the model parameters in two regression models with shared parameters. We assume sample sizes for the two groups are fixed and establish equivalence theorems to confirm the optimality of the design. As applications, we consider the parallel dose response model, the EMAX model and the Exponential model, each with shared parameters. The methodology is general and can be applied to other models or design problems. For example, we show the theoretical framework can be directly extended to the case when we are interested to find a c-optimal design to estimate the mean difference between the expected responses at an extrapolated dose for a nonlinear model, or when the total sample size for the whole study is fixed, and we wish to determine the optimal proportions of observations to allocate to the two groups, or we have multivariate responses. &lt;p&gt;Key words and phrases: Approximate design, equivalence theorem, group comparison, L-optimal design.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N310/j36N310.html</link>
<title> OPTIMAL AVERAGING ESTIMATION FOR DENSITY FUNCTIONS </title>
<author>Peng Lin, Jun Liao, Zudi Lu, Kang You and Guohua Zou </author>
<page>1235-1256. DOI:10.5705/ss.202022.0410</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Extraction of information from data is critical in the age of data science. Probability density function theoretically provides comprehensive information on the data. But, practically, different probability density models, either parametric or non-parametric, can often characterize partial features on the data, e.g., owing to model bias or less efficiency in estimation. In this paper we suggest a framework to optimally combine different density models to catch the comprehensive data features by a new information criterion (IC) based unsupervised learning approach. Our optimal information extraction is in the sense that the resultant density averaging or selected density minimises the Kullback&amp;ndash;Leibler (KL) information loss function. Differently from the usual supervised learning IC for model selection or averaging, we first need to derive an estimator of the KL loss function in our setting, which takes the Akaike and Takeuchi information criteria as two special cases. A feasible density model averaging (DMA) procedure is accordingly suggested, with the DMA estimation achieving the lowest possible KL loss asymptotically. Further, the consistency of the weights of the DMA estimator tending to the optimal averaging weights minimizing the KL distance is obtained, and the convergence rate of our empirical weights is also derived. Simulation studies show that the DMA performs overall better and more robustly than the commonly used parametric or nonparametric density models, including kernel, finite mixture, logarithmic scoring rule and selection methods for density estimation in the literature. The real data analysis further demonstrates the performance of the proposed method. &lt;p&gt;Key words and phrases: Asymptotic optimality, density averaging, density estimation, weight choice.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N311/j36N311.html</link>
<title> ASYMPTOTIC RESULTS FOR PENALIZED QUASI-LIKELIHOOD ESTIMATION IN GENERALIZED LINEAR MIXED MODELS </title>
<author>Xu Ning, Francis K.C. Hui and A.H. Welsh </author>
<page>1257-1278. DOI:10.5705/ss.202023.0343</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Generalized Linear Mixed Models (GLMMs) are widely used for analysing clustered data. One well-established method of overcoming the integral in the marginal likelihood function for GLMMs is penalized quasi-likelihood (PQL) estimation, although to date there are few asymptotic distribution results relating to PQL estimation for GLMMs in the literature. In this paper, we establish large-sample results for PQL estimators of the parameters and random effects in independent-cluster GLMMs, when both the number of clusters and the cluster sizes go to infinity. This is done under two distinct regimes: conditional on the random effects (essentially treating them as fixed effects) and unconditionally (treating the random effects as random). Under the conditional regime, we show the PQL estimators are asymptotically normal around the true fixed and random effects. Unconditionally, we prove that while the estimator of the fixed effects is asymptotically normally distributed, the correct asymptotic distribution of the so-called prediction gap of the random effects may in fact be a normal scale-mixture distribution under certain relative rates of growth. A simulation study is used to verify the finite-sample performance of our theoretical results. &lt;p&gt;Key words and phrases: Asymptotic independence, clustered data, large-sample distribution, longitudinal data, prediction.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N312/j36N312.html</link>
<title> GLOBAL GROUP TESTING AND SCREENING WITH DYNAMIC EFFECTS </title>
<author>Ying Cui and Limin Peng </author>
<page>1279-1300. DOI:10.5705/ss.202023.0285</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Identifying outcome-related variables is of general research interest in biomedical research. This task can be complicated by the presence of dynamic (or varying) variable effects that often manifest meaningful scientific mechanisms. Appropriately accounting for possible dynamic effects is crucial to avoid depreciating some important variables. In this work, we propose a model-free testing and screening framework by adopting a global view pertaining to the concept of interval quantile independence. The new framework not only permits robust identification of variables dynamically associated with an outcome, but also offers the flexibility to perform group testing that simultaneously evaluates multiple continuous or discrete covariates. We show that the key testing strategy can naturally evolve into unconditional and conditional screening procedures for ultra-high dimensional settings that enjoys the desirable sure screening property. We demonstrate good practical utility of the proposed methods via extensive simulation studies and a real application to a microarray data set. &lt;p&gt;Key words and phrases: Dynamic effects, hypothesis testing, interval quantile independence, variable screening.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N313/j36N313.html</link>
<title> TESTING FOR VARIANCE CHANGES UNDER VARYING MEAN AND SERIAL CORRELATION </title>
<author>Cheuk Wai Dominic Leung and Kin Wai Chan </author>
<page>1301-1322. DOI:10.5705/ss.202023.0238</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Detection of variance change points is statistically difficult when the data exhibit a varying mean structure and autocorrelation. Existing variance change point tests either require the assumption of mean constancy or sacrifice testing power due to serial dependence. This article addresses these problems by proposing a trend-robust and autocorrelation-efficient variance change point test via a differencing approach. This approach removes the mean effect without fitting the mean function. It also allows the test to retrieve the reduced power due to serial dependence. We prove that the optimal difference-based test should minimize the long-run coefficient of variation of the sample second moment of the noise instead of the long-run variance in the presence of serial dependence. The optimal solution can be efficiently computed by fractional quadratic programming. The asymptotic relative efficiency under a local alternative hypothesis is derived. A rate-optimal long-run variance estimator is also proposed. It is proven to be doubly robust against varying mean and variance change points. &lt;p&gt;Key words and phrases: Change point, cumulative sum, difference sequence, long-run variance, non-linear time series.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N314/j36N314.html</link>
<title> VALISE: A ROBUST VERTEX HUNTING ALGORITHM </title>
<author>Dieyi Chen, Zheng Tracy Ke and Shuyi Zhang </author>
<page>1323-1345. DOI:10.5705/ss.202023.0159</page>
</item>
<item>
<link>/statistica/J36N3/J36N315/j36N315.html</link>
<title> IDENTIFYING CAUSAL EFFECTS USING INSTRUMENTAL VARIABLES FROM THE AUXILIARY DATASET </title>
<author>Kang Shuai, Shanshan Luo, Wei Li and Yangbo He </author>
<page>1347-1366. DOI:10.5705/ss.202023.0088</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Instrumental variable approaches have gained popularity for estimating causal effects in the presence of unmeasured confounders. However, the availability of instrumental variables in the primary dataset is often challenged due to stringent and untestable assumptions. This paper presents a novel method to identify and estimate causal effects by utilizing instrumental variables from the auxiliary dataset, incorporating a structural equation model, even in scenarios with nonlinear treatment effects. Our approach involves using two datasets: one called the primary dataset with joint observations of treatment and outcome, and another auxiliary dataset providing information about the instrument and treatment. Our strategy differs from most existing methods by not depending on the simultaneous measurements of instrument and outcome. The central idea for identifying causal effects is to establish a valid substitute through the auxiliary dataset, addressing unmeasured confounders. This is achieved by developing a control function and projecting it onto the function space spanned by the treatment variable. We then propose a three-step estimator for estimating causal effects and derive its asymptotic results. We illustrate the proposed estimator through simulation studies, and the results demonstrate favorable performance. We also conduct a real data analysis to evaluate the causal effect between vitamin D status and body mass index. &lt;p&gt;Key words and phrases: Control function, data fusion, instrumental variable, unmeasured confounder.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N316/j36N316.html</link>
<title> ESTIMATION AND GOODNESS-OF-FIT TESTING FOR NON-NEGATIVE RANDOM VARIABLES WITH EXPLICIT LAPLACE TRANSFORM </title>
<author>Lucio Barabesi, Antonio Di Noia, Marzia Marcheselli, Caterina Pisani and Luca Pratelli </author>
<page>1367-1388. DOI:10.5705/ss.202023.0393</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Many flexible families of positive random variables exhibit non-closed forms of the density and distribution functions and this feature is considered unappealing for modelling purposes. However, such families are often characterized by a simple expression of the corresponding Laplace transform. Relying on the Laplace transform, we propose to carry out parameter estimation and goodnessof- fit testing for a general class of non-standard laws. We suggest a novel datadriven inferential technique, providing parameter estimators and goodness-of-fit tests, whose large-sample properties are derived. The implementation of the method is specifically considered for the positive stable and Tweedie distributions. A Monte Carlo study shows good finite-sample performance of the proposed technique for such laws. &lt;p&gt;Key words and phrases: Central limit theorem, consistent estimation, goodness-of-fit testing, Laplace transform, stable distribution, Tweedie distribution.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N317/j36N317.html</link>
<title> BALANCED SUBSAMPLING FOR BIG DATA WITH CATEGORICAL PREDICTORS </title>
<author>Lin Wang </author>
<page>1389-1406. DOI:10.5705/ss.202023.0434</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Supervised learning under measurement constraints is a common challenge in statistical and machine learning. In many applications, despite extensive design points, acquiring responses for all points is often impractical due to resource limitations. Subsampling algorithms offer a solution by selecting a subset from the design points for observing the response. Existing subsampling methods primarily assume numerical predictors, neglecting the prevalent occurrence of big data with categorical predictors across various disciplines. This paper proposes a novel balanced subsampling approach tailored for data with categorical predictors. A balanced subsample significantly reduces the cost of observing the response and possesses three desired merits. First, it is nonsingular and, therefore, allows linear regression with all dummy variables encoded from categorical predictors. Second, it offers optimal parameter estimation by minimizing the generalized variance of the estimated parameters. Third, it allows robust prediction in the sense of minimizing the worst-case prediction error. We demonstrate the superiority of balanced subsampling over existing methods through extensive simulation studies and a real-world application. &lt;p&gt;Key words and phrases: Data labeling, D-optimality, experimental design, orthogonal array, robust prediction.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N318/j36N318.html</link>
<title> POWERFUL SPATIAL MULTIPLE TESTING VIA BORROWING NEIGHBORING INFORMATION </title>
<author>Linsui Deng, Kejun He and Xianyang Zhang </author>
<page>1407-1433. DOI:10.5705/ss.202024.0152</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Clustered effects are often encountered in multiple hypothesis testing of spatial signals. In this paper, we propose a new method, termed two-dimensional spatial multiple testing (2D-SMT) procedure, to control the false discovery rate (FDR) and improve the detection power by exploiting the spatial information encoded in neighboring observations. The proposed method provides a novel perspective of utilizing spatial information by gathering signal patterns and spatial dependence into an auxiliary statistic. 2D-SMT rejects the null when a primary statistic at the location of interest and the auxiliary statistic constructed based on nearby observations are greater than their corresponding cutoffs. 2D-SMT can also be combined with different variants of the weighted BH procedures to improve the detection power further. A fast algorithm is developed to accelerate the search for optimal cutoffs in 2D-SMT. In theory, we establish the asymptotic FDR control of 2D-SMT under weak spatial dependence. Extensive numerical experiments demonstrate that the 2D-SMT method combined with various weighted BH procedures achieves the most competitive performance in FDR and power trade-off. &lt;p&gt;Key words and phrases: Empirical Bayes, false discovery rate, near epoch dependence, side information.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N319/j36N319.html</link>
<title> COMPLETE CONSECUTIVE ORDER-PAIRING DESIGN AND ITS DISTANCE-BASED LINEAR MODEL: DESIGN CONSTRUCTION AND ANALYSIS FOR ORDER-OF-ADDITION EXPERIMENTS </title>
<author>Jing-Wen Huang and Frederick Kin Hing Phoa </author>
<page>1435-1454. DOI:10.5705/ss.202023.0357</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; An order-of-addition (OofA) experiment investigates how the sequence of input factors influences the experimental response. This type of experiment has recently gain significant interest among practitioners in clinical trials and industrial processes. In this work, we introduce a new cost-efficient design called the Complete Consecutive Order-Pairing (CCOP) design. The CCOP design not only considers the effects of the component order on the response but also simultaneously accounts for the effects due to the component levels. We also propose a new statistical model associated with the CCOP design for identifying the optimal settings of both component order and levels. The CCOP design method evaluates the effects of two successive treatments by using the minimal number of runs, as each pair of level settings for two different components appears exactly once. Compared to recent studies on OofA experiments, our design effectively handles pure order experiments and multi-level experiments with a relatively small run size. &lt;p&gt;Key words and phrases: Clinical trials, cost-efficiency, order-of-addition experiments.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N320/j36N320.html</link>
<title> BAYESIAN OPTIMIZATION WITH PARETO-PRINCIPLED TRAINING FOR ECONOMICAL HYPERPARAMETER OPTIMIZATION </title>
<author>Yang Yang, Ke Deng and Yu Zhu </author>
<page>1455-1478. DOI:10.5705/ss.202023.0310</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; The specification of hyperparameters plays a critical role in determining the practical performance of a machine learning method. Hyperparameter Optimization (HPO), i.e., the searching for optimal specification of hyperparameters, however, often faces critical computational challenges due to the vast searching space and the high computational cost on model training under a given hyperparameter specification. In this paper, we propose BOPT-HPO, a systematic approach for efficient HPO by leveraging Bayesian optimization with Pareto-principled training, based on the observation that the training procedure of a machine learning method under a given hyperparameter specification often follows the Pareto principle (the 80/20 rule) that about 80% of the total improvement in the objective function is achieved in 20% of the training time. By introducing two levels of training corresponding to the Pareto principle, i.e., the eighty-percent training (ET) and the complete training (CT), and establishing a joint surrogate model for CT runs and ET runs, BOPT-HPO reduces the computational cost of HPO significantly under the framework of Bayesian optimization with multi-fidelity measurements. A wide range of experimental studies confirm that the proposed approach achieves economical HPO for various machine learning models, including support vector machines, feed-forward neural networks, and convolutional neural networks. &lt;p&gt;Key words and phrases: Automated artificial intelligence, black-box function optimization, computer experiments, multi-fidelity modelling, truncated Gaussian process.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N321/j36N321.html</link>
<title> OPTIMAL MODEL AVERAGING FOR IMBALANCED CLASSIFICATION </title>
<author>Ze Chen, Jun Liao, Wangli Xu and Yuhong Yang </author>
<page>1479-1498. DOI:10.5705/ss.202024.0012</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Imbalanced data with a high-dimensional input has been widely encountered in many areas of applications. In this situation, it usually becomes essential to reduce redundant variables via model selection to improve the classification performance. However, with a large number of variables, model selection uncertainty is typically very high. To deal with this problem, we present a feasible model averaging procedure based on a cost-sensitive support vector machine (CSSVM) coupled with a cost-sensitive data-driven weight choice criterion for imbalanced classification. Theoretical justifications are provided in two distinct scenarios. When the data exhibits a weak imbalance, we derive a relatively fast uniform convergence rate of the CSSVM solution. In contrast, when the data possesses a strong imbalance, the convergence rate becomes much slower. In both scenarios, an asymptotic optimality of the proposed model averaging approach in the sense of minimizing the out-of-sample hinge loss is established. Moreover, to reduce the computational burden imposed by a large number of candidate models for model averaging, we develop the CSSVM with an L&amp;lt;sub&amp;gt;1&amp;lt;/sub&amp;gt;-norm penalty to prepare candidate models. Numerical analysis shows the superiority of the proposed model averaging procedure over existing imbalanced classification methods. &lt;p&gt;Key words and phrases: Asymptotic optimality, imbalanced data, model averaging, uniform convergence rate.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N322/j36N322.html</link>
<title> ON DOUBLY ROBUST ESTIMATION WITH NONIGNORABLE MISSING DATA USING INSTRUMENTAL VARIABLES </title>
<author>Baoluo Sun, Wang Miao and Deshanee S. Wickramarachchi </author>
<page>1499-1519. DOI:10.5705/ss.202023.0383</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; Suppose we are interested in the mean of an outcome that is subject to nonignorable nonresponse. This paper develops new semiparametric estimation methods with instrumental variables which affect nonresponse, but not the outcome. The proposed estimators remain consistent and asymptotically normal even under partial model misspecifications for two variation independent nuisance components. We evaluate the performance of the proposed estimators via a simulation study, and apply them in adjusting for missing data induced by HIV testing refusal in the evaluation of HIV seroprevalence in Mochudi, Botswana, using interviewer experience as an instrumental variable. &lt;p&gt;Key words and phrases: Doubly robust estimation, endogeneous selection, exclusion restriction, instrumental variable, nonignorable nonresponse.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N323/j36N323.html</link>
<title> INFERENCE ON LARGE-SCALE GENERALIZED FUNCTIONAL LINEAR MODEL </title>
<author>Kaijie Xue and Riquan Zhang </author>
<page>1521-1540. DOI:10.5705/ss.202023.0356</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; In this work, we extend the classical generalized functional linear model to a large-scale generalized functional linear model to handle a variety of complex situations where the response (possibly discrete) can be nonlinearly linked to an ultra-high number of functional predictors. Unlike most existing requirements on functional data, we don't need to impose any conditions regarding eigenvalue-decay or square-integrability on those functional predictors, resulting in a more flexible but challenging model framework. Based on a penalized model estimator, we develop a general inferential method to assess the significance of an arbitrary group of regression curves. Concretely, a pseudo score function is adopted to construct the associated confidence region for the regression curves of interest. Notably, the proposed test is justified uniformly convergent to nominal level, without any demand on estimation consistency of the regression curves. Finally, numerical studies are carried out to show the empirical performance of the proposed test. &lt;p&gt;Key words and phrases: Eigenvalue-decay-free, estimation-consistency-relaxed, high dimensions, multiplier bootstrap, square-integrable-free.&lt;/span&gt;</description>
</item>
<item>
<link>/statistica/J36N3/J36N324/j36N324.html</link>
<title> FUNCTIONAL LINEAR MODELS WITH LATENT FACTORS </title>
<author>Zixuan Han, Tao Li, Jinhong You and Jiguo Cao </author>
<page>1541-1561. DOI:10.5705/ss.202024.0028</page>
<description>&lt;span style='font-size=12pt;'&gt;&lt;center&gt;Abstract&lt;/center&gt; We propose a novel functional linear model incorporating latent factors, where scalar response, scalar covariates, and functional covariates have repeated measurements for each subject. Our model accounts for latent factors that may impact the response but remain unobservable. To unveil and estimate these latent factors, we propose an iterated profile estimation method. We then establish the consistency and asymptotic properties of the estimators. To demonstrate the efficacy of our proposed estimation procedure, we conduct simulation studies across various scenarios. We compare our results with estimations derived from conventional functional linear models, revealing the superior performance of our method in addressing latent factors. We further illustrate our proposed model and methodology by analyzing real data from both financial markets and air pollution datasets. In these analyses, we successfully uncover hidden factors that exert influence in these specific fields. &lt;p&gt;Key words and phrases: Factor model, functional data analysis, functional regression, penalized spline, profile estimation.&lt;/span&gt;</description>
</item>
</channel>
</rss>
