Skip to content
Preprint

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Aug 2026 · 0 citations · 42 references
Mathematics Computer Science

TL;DR

This work proposes Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data, and proves that the WF estimator achieves minimax optimality over distribution families with bounded covariance.

Abstract

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.

View source

Similar papers

Preprint Jul 2026

Learning the Center and Radius of Wasserstein Ambiguity Sets for Data-Driven Decision Making

A more flexible framework in which a predictive model determines the nominal distribution and a separate model estimates a data-dependent radius is developed, which treats calibration as a practical mechanism for reliable decision making rather than a universal guarantee of improved optimization performance.

Junjie Guo · 1 citation

Robust Chance-Constrained Optimization using a Continuous Parameter Space Wasserstein-2 Ambiguity Set of Gaussian Mixtures

A novel formulation of a Wasserstein-2 metric that uses the Bures-Wasserstein (BW) metric over probability measures with finite second moments is developed, which allows the worst-case distribution to endogenously determine both how many mixture components receive mass and where their means and covariances lie within a...

Shibshankar Dey, Sanjay Mehrotra · 0 citations
Jul 2026

Generative Distributionally Robust Optimization

Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibility and adversarial structure: methods that accept arbitrary samplers do not restrict worst-case laws to a generator family, while generator-parameterized adversaries rely on model...

Zi-Wei Zhang, Jonathan Yu-Meng Li, Zhihao Jin · 0 citations
Preprint Aug 2026

Statistical Properties of Robust Learning under Distributional Shifts

Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data. Robust learning frameworks such as Distributionally Robust Optimization (DRO) and Robust Satisficing (RS) aim to address this challenge, yet their finite-sample guarantees under such...

Zhiyi Li, Xiaojie Mao, Yunbei Xu et al. · 0 citations
Jul 2026

Semi-Supervised Conditional Diffusion via Label Augmentation

Conditional diffusion models have become a powerful and flexible framework for learning complex conditional distributions from labeled data. In practice, however, acquiring high-quality labels is costly and time-consuming, leaving large volumes of unlabeled data unused. To address this, we introduce label-augmented con...

Jin Su, Yuan Gao, Yong Zhou et al. · 0 citations
Preprint Jul 2026

Amortized Inference for Sampling Distributions Where the Bootstrap Fails

A neural network is trained on simulated datasets drawn from a prior over a distribution family, using single independent draws of the root T_n - T(F) scored by the pinball loss, a proper scoring rule whose population minimizer is the posterior-predictive law of the root.

Akash Deep · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.