Skip to content
Preprint

Bias Reduction for Local Polynomial Derivative Estimation

Aug 2026 · 0 citations · 17 references
Mathematics

Abstract

Local polynomial smoothing is commonly used in non-parametric regression, but local linear derivative estimation still has a bias of order $O(h^2)$. This paper proposes an iterative data sharpening method to reduce the bias of derivative estimates while retaining the simplicity of local linear fitting. The method is based on two expectation operators: $L_0$, acting on the regression function, and $L_1$, acting on the first-order derivative. By repeatedly applying the residual operator $R=I-L_0$, a series of sharpened derivative estimates can be constructed. After $l$ sharpening steps, the bias order can be reduced from $O(h^2)$ to $O(h^{2l+2})$. For the Gaussian kernel, all sharpening coefficients equal 1, giving a simple closed-form single-bandwidth expression. Simulation experiments on three smooth test functions show that this method can significantly reduce the estimation bias while revealing a bias-variance trade-off.

View source

Similar papers

Preprint Jul 2026

Exact Generalization Error Curves of Kernel Ridge Regression for Functional Moment Estimation

Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance estimation, and higher-order moment estimation in a single framework. Our main result gives a precise $1+o_{\mathbb{P}}(1)$ expansion for the $L^2$ error at each admissible regularization parameter. The expansion consists of bias and three variance terms corresponding respectively to variation across the independent sample paths, latent signal variation at each sample point, and variation from measurement errors, identifying the refined error structure underlying functional data. As applications, we show that KRR attains the minimax rate for source smoothness $s \leq 2$ but becomes suboptimal in the sparse regime for $s>2$ due to saturation. A technical ingredient is a set of concentration inequalities for $U$-statistics suited to the dependent product structure of functional observations.

Yi Ding, Yicheng Li · 0 citations
Preprint Aug 2026

A Simple Approximation to the Distribution of the Ridge Regression Estimator

We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.

J. M. Olea, Ryan Strong, Amilcar Velez et al. · 0 citations
#machine learning Preprint Aug 2026

Generalized Splines and Gaussian Processes

For finite-dimensional linear inverse problems where the variables are Gaussian, it is well-known that the minimum-mean-square error estimator takes the form of a regularized least-squares data fit. In this chapter, we show that this equivalence extends to a much broader infinite-dimensional setting where generalized splines take the role of linear regressors and generalized Gaussian processes on a nuclear space $S$ are the counterpart of Gaussian random vectors. The scope of this extension is of the same nature as the switch from the classic notion of function to that of a distribution, also known as a"generalized function."Our formalism involves a whitening/regularization operator $L: S\to S'$ whose continuous extension induces a native Hilbert space $H\subset S'$ that plays a central role in our characterization. The presentation is self-contained for the most part and remarkably general and powerful. It allows for the recovery of all known instances of such equivalences; in particular, the methods involving innovations and reproducing-kernel Hilbert spaces developed by Kailath and his students, and the mathematical correspondence between fractional splines and Mandelbrot's fractional Brownian motion (fractals), with the former being the optimal estimators of the latter. It also covers general Bayesian methods for the resolution of infinite-dimensional inverse problems.

Michael Unser · 0 citations
Preprint Aug 2026

Sharp proper estimation of fixed-component Gaussian location mixtures in polynomial time

Exhaustive moment fitting in this constant-dimensional space produces a proper mixture and, together with the dimension-free moment characterization of Gaussian mixtures, achieves the optimal Hellinger rate in polynomial arithmetic time for every fixed $k$.

Heng-Zhi He, Guang Cheng · 0 citations
Preprint Aug 2026

On the Iterate Convergence of AdaGrad for Generalized Smooth Convex Optimization

This work constructs a counterexample empirically showing that smoothness alone is not sufficient for the sequential convergence of AdaGrad-type algorithms, and suggesting that additional geometric hypotheses are indispensable for sequential convergence results.

Mathieu Besançon, Tung Le · 0 citations
Preprint Aug 2026

Generalization Error Estimation for Primal--Dual Algorithms in Non-Smooth Regression

This paper studies trajectory-wise estimation of generalization error for primal--dual algorithms in non-smooth regression. Motivating examples include \(\ell_1\)-penalized least absolute deviations regression and square-root Lasso regression, where the data-fitting loss is non-differentiable and existing risk estimators for gradient-type optimization paths do not apply directly. We develop a general recursive framework that includes the Chambolle--Pock algorithm and related primal--dual splitting methods. We estimate risk by correcting each in-sample fitted value with a weighted combination of past dual iterates. The ideal weights are Stein derivative contractions and depend on the design covariance. We construct replacement weights from observable derivative contractions of the fitted-signal trajectory, yielding a covariance-free, data-driven correction. For high-dimensional Gaussian designs and fixed finite iteration horizon, we prove finite-sample guarantees for both estimators. For square-root ridge, we further establish a matched-Gaussian universality result beyond Gaussian designs. Numerical experiments show that the proposed estimators accurately track the out-of-sample risk along finite optimization paths.

Kai Tan, Pierre C. Bellec · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.