High-dimensional change-point analysis is essential in modern statistical inference. However, existing methods are often designed either for specific parameters (e.g., mean or variance) or for particular tasks (e.g., testing or estimation), making them difficult to generalize. Moreover, they typically rely on restrictive distributional assumptions, limiting their robustness to heavy-tailed data. We propose a unified framework for testing, estimating, and inferring multiple change points in high-dimensional data. Our approach leverages a two-sample U-statistic within a moving window, allowing flexible kernel function selection to accommodate structural changes in general parameters such as variance changes or robust statistics. For testing, we develop an L-infinity norm-based statistic with a high-dimensional multiplier bootstrap procedure, achieving minimax-optimal power under sparse alternatives. For estimation, we construct an initial estimator for the change-point number and locations and refine it using the U-statistic Projection Refinement Algorithm (U-PRA), attaining minimax-optimal localization rates. We further derive the asymptotic distribution of refined estimators, enabling valid confidence interval construction. Extensive numerical experiments demonstrate the better performance of our method across various settings, including heavy-tailed distributions. Applications to genomic copy number variation data highlight its practical utility. An R package implementing the proposed method, U-PRA, is publicly available at https://github.com/liubin0145/R-codes-UPRA/.
We develop a framework for simultaneous change-point inference of high-dimensional functional time series. The observations are modeled as temporally dependent vectors whose coordinates take values in possibly different separable Hilbert spaces, thereby covering a broad class of functional data. Heterogeneous mean changes may occur at coordinate-specific locations, and the contemporaneous dependence across coordinates is left unrestricted. Our procedure is based on coordinatewise cumulative-sum statistics and a residual block multiplier bootstrap that provides a common critical value for the global test and all coordinatewise decisions. We establish a nonasymptotic Gaussian approximation, quantitative nonasymptotic bounds for strong family-wise error control under arbitrary mixtures of changed and unchanged coordinates, covariance-adaptive detection guarantees, and simultaneous high-probability bounds for change-point localization. The bounds accommodate high-dimensional regimes in which the number of functional coordinates grows exponentially in a power of the sample size. We investigate finite-sample performance in simulations and illustrate the method using river discharge curves and high-frequency financial log returns.
We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.
Partial missingness is common in high-dimensional data, but most existing change-point procedures are developed for fully observed sequences. We introduce gMiss, a graph-based framework for testing and localizing a change in the observed-data distribution of a partially observed high-dimensional sequence. The method treats the observed values together with the missingness indicators as the object of inference, so the target alternative is a change in the induced observed data law. It is designed for general distributional changes and requires neither sparsity nor Gaussianity. When the augmented observations are independent, the full permutation test controls type I error in finite samples. The procedure combines graph scans based on elementwise imputation and distance imputation. The two scans capture complementary graph patterns. Simulation results indicate that gMiss maintains accurate null calibration across the MCAR and MAR designs considered, remains competitive under Gaussian location alternatives, and exhibits strong power and localization performance in many non-Gaussian location and scale settings. We further illustrate the practical utility of the method through an application to genomic copy-number data, where gMiss identifies additional candidate boundaries that are visually plausible in the raw heatmap.
Modern time series are often long, serially dependent, and non-stationary. Existing change-point methods either target specific changes or become computationally intensive when using nonparametric costs on long series. Many also require thresholds to be carefully calibrated under serial dependence. We introduce SCAN, an offline method for detecting multiple distributional change-points in long, serially dependent univariate time series. SCAN compares adjacent windows using an integral probability metric, calibrates local discrepancies with a dependence-aware bootstrap, and refines candidate locations using a scaled 1-Wasserstein criterion, enabling detection of changes in mean, variance, and broader distributional structure within a unified framework. An ensemble over multiple window sizes reduces sensitivity to window size and threshold specification. We establish consistency of the estimated number and locations of change-points under exponential alpha-mixing dependence, and show that the localization statistic reduces to a CUSUM-type statistic under pure mean shifts. In simulations with up to one million observations, SCAN generally achieves higher covering and F1-scores than competing methods across mean and joint mean-variance shifts, particularly under serial dependence. On real data, SCAN identifies labeled activity transitions in sensor data and interpretable structural changes in hourly Bitcoin prices. Implementations are available in the Python package scan-cpd and R package scanr.
Ashoka Prabashwara, P. Menéndez, Liam Hodgkinson et al.· 0 citations
Classical MANOVA procedures are not directly applicable in high-dimensional settings where the number of variables is comparable to, or exceeds, the sample size, and many existing high-dimensional MANOVA tests remain sensitive to outlying observations. This study proposes a weighted minimum regularized covariance determinant (MRCD)-based robust Wilks’ Lambda test for one-way high-dimensional MANOVA. The proposed method combines MRCD-based robust location and scatter estimation with a robust distance-based reweighting step and uses permutation calibration to obtain p-values. Through extensive Monte Carlo simulations, the method is evaluated in terms of Type-I error control, power, and robustness under structured contamination. Under clean data, the proposed test maintains empirical Type-I error near the nominal level, with only modest aggregate differences from Cheng-GM; Schott’s test can have higher power under weak signals. Under contaminated null scenarios where outliers create artificial group separation, the proposed method yields lower false-rejection rates than the competitors considered. A controlled sensitivity illustration using breast-cancer gene-expression data shows the same qualitative behavior after imposed contamination. The method is therefore positioned as a robustness-oriented option for contamination-prone high-dimensional MANOVA, at the cost of additional computation.
We propose a unified ridge-regularized Hotelling framework for detecting and locating mean changes in functional time series. A growing basis expansion converts the functional observations into high-dimensional score vectors. Their long-run covariance is estimated by an edge-corrected difference-based procedure. Ridge regularization stabilizes inference under spectral decay. An explicit local-power formula shows that the power-maximizing ridge depends on the unknown spectral orientation of the change. We therefore combine a family of ridge CUSUM statistics by a Cauchy transform and calibrate the aggregate directly from their joint weighted-bridge limit. For multiple changes, we embed local maximum-ridge statistics in a wild binary segmentation procedure, followed by local refinement. Under mild conditions, we establish the validity, local power, consistency, and localization properties of the proposed tests. In the multiple-change setting, the procedure consistently recovers the number of changes and uniformly estimates their locations. The framework accommodates weak dependence and non-Gaussian functional errors. Simulations and two empirical applications demonstrate the favorable finite-sample performance of the proposed methods.
Ping Zhao, Long Feng· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.