This manuscript develops a non-parametric and robust framework for estimating the scale of additive noise in weakly sparse systems. The method does not require independence, prescribed dependence, or temporal regularity of the noise sequence. We introduce a class of order-statistic estimators based on comparing the sorted observations with deterministic or random proxies generated from a reference noise distribution. This purely spatial approach avoids preliminary filtering or temporal decorrelation, and therefore preserves the sparsity structure of the latent signal. We establish non-asymptotic concentration inequalities for weighted loss functions, with bounds that separate the contribution of the signal from the discrepancy between the ordered noise and the proxy. We then control this proxy discrepancy in independent and correlated regimes, including heavy-tailed reference laws. Finally, we apply the method to high-frequency observations of continuous-time stochastic processes, obtaining scale estimators for fractional Brownian motion and stable L\'evy noise in the presence of lower-variation additive perturbations.
We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.
This paper develops a unified asymptotic theory for inverse-probability-weighted conditional U-statistics of arbitrary fixed order in the presence of missing-at-random responses and infinite-dimensional functional covariates. The target is a conditional higher-order functional generated by a measurable response kernel and evaluated locally on a separable Banach space. Localization is formulated through delta sequences, providing a common framework for kernel, partition, regressogram, orthogonal series, and related smoothing procedures without recourse to finite-dimensional density arguments. For bounded kernels, we establish uniform almost-complete convergence over pseudo-compact functional domains and obtain a sharp decomposition into deterministic localization bias and stochastic fluctuation. The latter is governed by the localized-kernel variance, the envelope of the delta sequence, the metric complexity of the indexing domain, and the small-ball concentration of the functional covariate. Unbounded kernels are treated under explicit weighted moment, truncation, and summability conditions. The feasible theory quantifies the additional perturbation induced by estimating the propensity score and identifies conditions under which this first-stage uncertainty is asymptotically negligible. Pointwise distributional theory is derived through a denominator linearization combined with the Hoeffding decomposition of the centered localized kernel. The Gaussian limit is driven by the first projection, while the higher-order canonical components are shown to be negligible under explicit local-mass, moment, and noncancellation assumptions. This yields oracle-equivalent feasible inference, a consistent first-projection variance estimator, and asymptotically valid studentized confidence intervals. A finite-grid adaptive comparison principle is also developed for data-driven resolution selection. The scope of the theory is illustrated through conditional rank functionals, discrimination with incomplete labels, metric-learning criteria, and functional prediction. Synthetic and semi-synthetic studies based on functional classification, phoneme log-periodograms, and growth trajectories document the finite-sample interaction between covariate-dependent label observation, local information loss, propensity estimation, and inverse-weighting variance.
We study estimation and inference for a semiparametric class of time series models that specify only the conditional expectation, which is a known link function applied to a linear combination of past observations and covariates. The class covers count, binary, bounded and conditionally heteroskedastic responses within a single formulation, and the parameter is estimated by a quasi-likelihood estimating equation based on the first conditional moment. Under stationarity and a weak-dependence condition expressed through the functional dependence measure, we establish two results. First, using a Fuk--Nagaev inequality for weakly dependent sequences, we show that the estimator is localized in a shrinking neighbourhood of the true value with probability $1-o(n^{-1/2})$. Second, combining a Berry--Esseen bound for weakly dependent sequences with a Gaussian anti-concentration argument to control the remainder of the linear expansion, we obtain a Berry--Esseen bound for linear projections of the estimator, uniform over projection directions. From the projected bound we derive studentized confidence intervals with explicit coverage error and a conservative Bonferroni test for linear hypotheses on the parameters. For real data analysis, we extend the Beta autoregression for double-bounded data to an arbitrary link given by the inverse of a distribution function, and apply it to ten pairwise realized correlations of large-cap technology-stock returns, using Nasdaq and Dow~Jones index returns as covariates.
In this paper, we study the autocovariance matrix estimation and inference problems under heavy-tailedness, high-dimensionality, general nonlinear temporal dependence, and potentially nonstationarity of time series. We consider two types of tail-robust autocovariance matrix estimation methods: the element-wise Huber's $M$-estimator and a computationally more efficient element-wise truncated estimator. Both estimators are designed to achieve sharp error bounds in matrix max-norm. The nonasymptotic properties of these estimators are proved based on new variants of Bernstein-type inequalities under functional dependence for the potentially nonstationary processes which may be of independent interest. Moreover, we prove a high-dimensional Gaussian approximation result, as a limiting distribution, for our element-wise truncated autocovariance estimator. A Gaussian multiplier bootstrap result is also given to facilitate the practicality. Our theoretical results are nonasymptotic, which gives explicit error bounds in terms of the sample size, dimensionality, moments, and the strength of temporal dependence. Numerical evidence is provided to support our theoretical results. Finally, we illustrate the benefits of the proposed methodology for detecting change points in monthly macroeconomic data.
Hao-Tian Xu, S. Guerrier, Run-Ze Li et al.· 0 citations
In the era of high-dimensional data, the classical assumption that the number of observations n vastly exceeds the number of variables p is frequently violated. When p and n grow proportionally (p/n → c > 0), the sample covariance matrix becomes severely distorted by sampling noise. Its eigenvalues are systematically biased: large population variances are overestimated, and small ones are underestimated. This phenomenon, governed by the Marchenko–Pastur law of Random Matrix Theory (RMT), renders standard statistical procedures highly unstable. This paper provides a comprehensive, mathematically rigorous treatment of spectral shrinkage, the optimal remedy for this distortion. We transition from the theoretical foundations of the Stieltjes transform to the practical implementation of rotationally invariant estimators. By combining formal proofs, geometrical interpretations, and reproducible R simulations with explicit console outputs, we demonstrate why spectral shrinkage is not merely a heuristic regularization technique, but a mathematically undeniable necessity for modern high-dimensional statistics.
Innocent Nsabimana· International Journal For Mu...· 0 citations
We develop a design-conditional limit theory for kernel estimators of conditional U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed to take values in a general Polish space, and the target is indexed by a class of symmetric kernels of a fixed order. Functional localization is induced by single-index semi-metrics, while spatial localization is performed on the rescaled observation domain. Missing responses are incorporated through a complete-case construction under a Missing At Random condition and a uniform-positivity assumption. The resulting estimator is a ratio of spatially weighted U-statistics with random tuplewise observation indicators. The asymptotic analysis must account simultaneously for four sources of complexity: dependence within the spatial field, nonstationarity across an expanding domain, concentration in an infinite-dimensional covariate space, and the random thinning generated by missing responses. Conditioning on the sampling locations removes the randomness of the spatial design weights but does not eliminate dependence among the observations. We therefore derive a design-conditional projection decomposition adapted to the triangular-array structure of the model. The leading component is represented by a spatially dependent complete-case empirical process, whereas the higher-order canonical terms are controlled uniformly over the response kernels, functional-target points, single-index directions, and rescaled spatial locations. The proofs combine stationary tangent-field approximations for locally stationary random fields, large-block–small-block decompositions, coupling arguments under spatial absolute regularity, small-ball probability estimates, and entropy bounds for the joint indexing class. These arguments yield a uniform stochastic expansion in which the empirical fluctuation, the spatial–functional smoothing bias, and the local-stationarity approximation error appear as distinct contributions. In particular, the local-stationarity remainder has no counterpart in the strictly stationary theory and quantifies the cost of replacing the observed nonstationary field with its stationary tangent approximation. Under the MAR and positivity conditions, complete-case sampling reduces the effective local information and modifies the covariance structure, but it does not change the formal order of the uniform-convergence rate. Under strengthened moment, mixing, entropy, and negligibility conditions, we establish weak convergence of the normalized conditional U-process in the corresponding supremum-norm function space to a tight centered Gaussian process. The limiting covariance is determined by the complete-case first-order projection and consequently retains the effect of the observation propensity and the spatial dependence structure. We also introduce a complete-case leave-tuple-out spatial prediction criterion for bandwidth selection and prove oracle optimality over admissible bandwidth families. The general theory applies to conditional rank association, discrimination probabilities, set-indexed conditional distribution functionals, and related pairwise statistical-learning criteria. Simulation experiments and applications to spatial environmental and epidemiological data illustrate the finite-sample implications of the theory and the stabilizing role of single-index localization. Viewed through the lens of data-driven science, the framework addresses a fundamental asymmetry between the information carried by irregular, locally heterogeneous functional covariates and the selectively observed response tuples. By combining design conditioning, complete-case normalization, tangent-field localization, and single-index dimension reduction, the proposed approach resolves this inferential asymmetry at the level of the model by matching estimation and uncertainty quantification to the information actually available locally, without imposing artificial stationarity or complete-data symmetry.
Salim Bouzebda· Symmetry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.