Skip to content

Author

Salim Bouzebda

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Learning Nonparametric Conditional Single-Index U-Processes for Missing Locally Stationary Functional Random Fields with Stochastic Spatial Design

We develop a design-conditional limit theory for kernel estimators of conditional U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed to take values in a general Polish space, and the target is indexed by a class of symmetric kernels of a fixed order. Functional localization is induced by single-index semi-metrics, while spatial localization is performed on the rescaled observation domain. Missing responses are incorporated through a complete-case construction under a Missing At Random condition and a uniform-positivity assumption. The resulting estimator is a ratio of spatially weighted U-statistics with random tuplewise observation indicators. The asymptotic analysis must account simultaneously for four sources of complexity: dependence within the spatial field, nonstationarity across an expanding domain, concentration in an infinite-dimensional covariate space, and the random thinning generated by missing responses. Conditioning on the sampling locations removes the randomness of the spatial design weights but does not eliminate dependence among the observations. We therefore derive a design-conditional projection decomposition adapted to the triangular-array structure of the model. The leading component is represented by a spatially dependent complete-case empirical process, whereas the higher-order canonical terms are controlled uniformly over the response kernels, functional-target points, single-index directions, and rescaled spatial locations. The proofs combine stationary tangent-field approximations for locally stationary random fields, large-block–small-block decompositions, coupling arguments under spatial absolute regularity, small-ball probability estimates, and entropy bounds for the joint indexing class. These arguments yield a uniform stochastic expansion in which the empirical fluctuation, the spatial–functional smoothing bias, and the local-stationarity approximation error appear as distinct contributions. In particular, the local-stationarity remainder has no counterpart in the strictly stationary theory and quantifies the cost of replacing the observed nonstationary field with its stationary tangent approximation. Under the MAR and positivity conditions, complete-case sampling reduces the effective local information and modifies the covariance structure, but it does not change the formal order of the uniform-convergence rate. Under strengthened moment, mixing, entropy, and negligibility conditions, we establish weak convergence of the normalized conditional U-process in the corresponding supremum-norm function space to a tight centered Gaussian process. The limiting covariance is determined by the complete-case first-order projection and consequently retains the effect of the observation propensity and the spatial dependence structure. We also introduce a complete-case leave-tuple-out spatial prediction criterion for bandwidth selection and prove oracle optimality over admissible bandwidth families. The general theory applies to conditional rank association, discrimination probabilities, set-indexed conditional distribution functionals, and related pairwise statistical-learning criteria. Simulation experiments and applications to spatial environmental and epidemiological data illustrate the finite-sample implications of the theory and the stabilizing role of single-index localization. Viewed through the lens of data-driven science, the framework addresses a fundamental asymmetry between the information carried by irregular, locally heterogeneous functional covariates and the selectively observed response tuples. By combining design conditioning, complete-case normalization, tangent-field localization, and single-index dimension reduction, the proposed approach resolves this inferential asymmetry at the level of the model by matching estimation and uncertainty quantification to the information actually available locally, without imposing artificial stationarity or complete-data symmetry.

Salim Bouzebda · 0 citations
Open access Aug 2026

Asymptotic Normality of Wavelet Density and Regression Estimators Under Censored Ergodic Observations

This paper develops a pointwise distributional theory for linear wavelet density and regression estimation from randomly right-censored observations exhibiting stationary ergodic dependence. In contrast to the prevailing literature, which typically relies on quantitative mixing conditions, our analysis is conducted under ergodicity alone, thereby encompassing substantially broader classes of dependent processes. We establish asymptotic normality for an oracle inverse-probability-weighted estimator based on the true censoring distribution and for its feasible counterpart obtained through Kaplan–Meier substitution. A central result shows that estimating the censoring distribution has no first-order effect on the limiting law, so that the feasible and oracle procedures are asymptotically equivalent. The proof strategy departs from conventional covariance inequalities and blocking arguments and instead combines a martingale-predictable decomposition with martingale central limit theory and ergodic convergence of conditional moments. The framework is further extended to a broad family of wavelet regression functionals involving transformed responses. To render the asymptotic theory directly usable for statistical inference, we introduce a randomly weighted procedure that consistently reproduces the limiting distribution of the feasible estimator. This yields asymptotically valid pointwise confidence intervals without requiring explicit estimation of the unknown asymptotic variance or the introduction of additional smoothing parameters. The scope of the theory includes several important non-mixing and long-range dependent models, while an extensive simulation study demonstrates the finite-sample accuracy and robustness of the proposed inferential methodology.

Salim Bouzebda, S. Didi · 0 citations
Open access Aug 2026

Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling

We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.

Salim Bouzebda, S. Didi · 0 citations
Open access Aug 2026

Statistical Learning Theory for Inverse-Probability-Weighted Conditional U-Statistics via Delta Sequences Under Functional Missing-at-Random Models

This paper develops a unified asymptotic theory for inverse-probability-weighted conditional U-statistics of arbitrary fixed order in the presence of missing-at-random responses and infinite-dimensional functional covariates. The target is a conditional higher-order functional generated by a measurable response kernel and evaluated locally on a separable Banach space. Localization is formulated through delta sequences, providing a common framework for kernel, partition, regressogram, orthogonal series, and related smoothing procedures without recourse to finite-dimensional density arguments. For bounded kernels, we establish uniform almost-complete convergence over pseudo-compact functional domains and obtain a sharp decomposition into deterministic localization bias and stochastic fluctuation. The latter is governed by the localized-kernel variance, the envelope of the delta sequence, the metric complexity of the indexing domain, and the small-ball concentration of the functional covariate. Unbounded kernels are treated under explicit weighted moment, truncation, and summability conditions. The feasible theory quantifies the additional perturbation induced by estimating the propensity score and identifies conditions under which this first-stage uncertainty is asymptotically negligible. Pointwise distributional theory is derived through a denominator linearization combined with the Hoeffding decomposition of the centered localized kernel. The Gaussian limit is driven by the first projection, while the higher-order canonical components are shown to be negligible under explicit local-mass, moment, and noncancellation assumptions. This yields oracle-equivalent feasible inference, a consistent first-projection variance estimator, and asymptotically valid studentized confidence intervals. A finite-grid adaptive comparison principle is also developed for data-driven resolution selection. The scope of the theory is illustrated through conditional rank functionals, discrimination with incomplete labels, metric-learning criteria, and functional prediction. Synthetic and semi-synthetic studies based on functional classification, phoneme log-periodograms, and growth trajectories document the finite-sample interaction between covariate-dependent label observation, local information loss, propensity estimation, and inverse-weighting variance.

Salim Bouzebda · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.