Skip to content
Preprint

Distributed Stochastic Smoothing ADMM for Penalized Quantile Regression

Aug 2026 · 0 citations · 33 references
Mathematics

TL;DR

A distributed stochastic smoothing alternating direction method of multipliers (DSS-ADMM) for horizontally partitioned penalized quantile regression, which characterize the scope of an extension to the minimax concave penalty and the smoothly clipped absolute deviation penalty.

Abstract

Quantile regression is well suited to heterogeneous and heavy-tailed data, but computation becomes challenging for large, distributed data sets because the check loss is nonsmooth. We propose a distributed stochastic smoothing alternating direction method of multipliers (DSS-ADMM) for horizontally partitioned penalized quantile regression. Each worker computes a mini-batch gradient of a Huber-smoothed check loss, and a coordinator performs a single proximal aggregation step for the regularizer. Raw observations remain local, worker updates run in parallel, and the method requires no matrix inversion. For proper, closed, and convex penalties, stacking the local coefficient vectors yields a standard two-block stochastic ADMM formulation. With fixed smoothing, we establish an expected $O(\log K/\sqrt K)$ joint objective-feasibility bound and an explicit $\eps/4$ approximation term for the original check-loss objective; when the smooth block is strongly convex, the bound improves to $O(\log K/K)$. We also characterize the scope of an extension to the minimax concave penalty and the smoothly clipped absolute deviation penalty. Reproducible simulations consider both homogeneous worker partitions, in which observations are independently and identically distributed across workers, and heterogeneous partitions, in which worker-specific covariate distributions differ. Sensitivity studies and analyses of the diabetes and Engel data illustrate the trade-offs among per-observation gradient evaluations, communication, consensus, sparsity, and prediction.

View source

Similar papers

Jul 2026

Tensor Quantile Regression With Exponential‐Type Penalty

A robust tensor quantile regression method, in which CANDECOMP/PARAFAC (CP) decomposition is employed for dimension reduction, and an exponential‐type penalty (ETP) is imposed at the element‐wise level to achieve sparse variable selection.

Tan Meng, Shuo Liu, Mao-Zai Tian · 0 citations
Jul 2026

Adaptive deep nonparametric regression from dependent data under covariate shift

Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions. In this case, the standard metric under the source distribution is not appropriate. This paper considers deep neural network estimators for nonparametric quantile and Huber regression under covariate shift and from dependent observations. We deal with a generalized Bernstein-type inequality that is satisfied by many classical models, including i.i.d. observations, $\phi$-mixing, strong mixing, and $\mathcal{C}$-mixing processes. To perform the covariate shift phenomenon, we propose a sparse-penalized deep neural network (SPDNN) estimator that takes into account the discrepancy between the source and target distributions of the data. When the density ratio (between the source and target distributions of the covariate) is unknown, a two steps pre-training procedure is carried out: the first step is devoted to the construction of a least squares SPDNN estimator of the density ratio; which is used in the second step to perform a pre-training reweighted SPDNN estimator of the regression function. For both the quantile and the Huber regression, non-asymptotic error bounds of the proposed SPDNN estimators are established in the class of H\"older smooth functions. These estimators can adaptively attain (up to a logarithmic factor) the minimax optimal convergence rate from i.i.d. data as well as from several classical time series models.

W. Kengne, Ehud Mossa Ockegna · 0 citations
Preprint Aug 2026

Generalization Error Estimation for Primal--Dual Algorithms in Non-Smooth Regression

This paper studies trajectory-wise estimation of generalization error for primal--dual algorithms in non-smooth regression. Motivating examples include \(\ell_1\)-penalized least absolute deviations regression and square-root Lasso regression, where the data-fitting loss is non-differentiable and existing risk estimators for gradient-type optimization paths do not apply directly. We develop a general recursive framework that includes the Chambolle--Pock algorithm and related primal--dual splitting methods. We estimate risk by correcting each in-sample fitted value with a weighted combination of past dual iterates. The ideal weights are Stein derivative contractions and depend on the design covariance. We construct replacement weights from observable derivative contractions of the fitted-signal trajectory, yielding a covariance-free, data-driven correction. For high-dimensional Gaussian designs and fixed finite iteration horizon, we prove finite-sample guarantees for both estimators. For square-root ridge, we further establish a matched-Gaussian universality result beyond Gaussian designs. Numerical experiments show that the proposed estimators accurately track the out-of-sample risk along finite optimization paths.

Kai Tan, Pierre C. Bellec · 0 citations
Preprint Aug 2026

Variable Smoothing for Weakly Convex Problems with Non-Euclidean Directions

We propose MELMO (Moreau Envelope Smoothing with Linear Minimization Oracles), an algorithm for composite optimization problems of the form min x f (x) + g(T x), where f is smooth and g may be non-smooth. The method leverages the Moreau envelope to smooth the non-smooth component while adapting to problem geometry through linear minimization oracles. Assuming g is $\rho$-weakly convex, we establish a family of convergence bounds parameterized by the step-size and smoothing schedules, thereby making explicit the trade-off between optimizing the smoothed objective and recovering stationarity for the original composite problem. In particular, one regime yields O(k -1/4 ) rates for both the smoothed-gradient norm and a composite stationarity proxy, while another yields O(k -1/3 ) for the smoothed-gradient norm together with O(k -1/4 ) for the composite proxy. We also establish a K-horizon-dependent convergence rate that yields O(K -1/3 ) for the composite proxy. Empirically, MELMO is competitive with variable smoothing and subgradient baselines on sparse low-rank matrix factorization and image denoising.

Farid Najar · 0 citations
Preprint Jul 2026

Finite-horizon quantile martingale posteriors: raw-urn laws and matrix-gain regression

Martingale posteriors quantify uncertainty by forward-imputing observations from one-step-ahead predictive distributions, but implementations stop after finitely many imputations. For the empirical P\'olya-urn posterior of a quantile the law of the stopped state is derived. The quantile of the stopped urn measure keeps the familiar martingale tail-sum variance fraction; the deployed stochastic-approximation tracker with frozen gain $c$ does not. Its variance carries an explicit factor $G_a$ with $a=cf_0(q_\tau)$, which may fall below or exceed the tail fraction, and a density-adapted gain restores calibration through a density-free inflation. Shared urn innovations yield the joint law of finitely many quantile levels. For conditional quantile regression, a smoothed martingale posterior started at the ordinary quantile-regression estimator with a full inverse-Jacobian matrix gain satisfies a process Bernstein--von Mises theorem with calibrated finite-horizon bands; scalar or diagonal gains cannot match the sandwich covariance process.

N. Le · 0 citations
Preprint Aug 2026

Distributed Selective Inference for Quantile Regression

We propose a distributed selective inference framework tailored for high-dimensional quantile regression. To enable valid post-selection inference in this context, we address the computational challenge posed by the non-smooth quantile loss via a response-surrogation strategy. This strategy transforms the problem into a penalized least-squares formulation, thereby facilitating the application of distributed selective inference. For valid post-selection inference, a randomized procedure is introduced, in which the Lasso selection event is characterized through the associated Karush-Kuhn-Tucker conditions and the conditional distribution of the aggregated estimator is derived given the selection event. The resulting algorithm requires only three rounds of communication between local machines and the central server. Under standard regularity conditions, we establish the asymptotic validity of the proposed procedure and develop a large-deviation approximation to the selective likelihood for computationally tractable implementation. Simulation studies and a real-data application demonstrate the satisfactory finite-sample performance of the proposed method.

Xiaohui Yuan, Jiahan Teng, Yan Zhou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.