Skip to content

Unifying Distributional Training for One-Step Visual Generation

Sep 2026 · 2 citations · 74 references
Computer Science

TL;DR

A unified theoretical framework is introduced that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow and motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations.

Abstract

Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with 1.45 $\mathrm{FDr}^6$ on pMF-H and 1.64 on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore.

View source

Similar papers

#machine learning Preprint Sep 2026

Data Unlearning via Inverse Distillation

Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step...

Aleksei Leonov, N. Kornilov, Zhen-He Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Learning a Flow to Self-Supervised Representations

Flow-Based Distribution Matching is introduced, a non-adversarial framework that learns this reference-directed geometry through spherical conditional velocity regression that achieves performance nearly on par with DM and remains competitive with existing SSL methods.

Yu-Ling Jiao, Wen-Sen Ma, Hou-Duo Qi et al. · 0 citations
Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
Preprint Sep 2026

Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...

Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al. · 0 citations
Preprint Aug 2026

Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges

This work introduces a novel framework, Gaussian Bridge Consistency (GBC), to address challenges of semi-supervised learning by constructing semantic interpolation paths between unlabeled samples and high-quality class anchors, and proposes BridgeMix, a confidence-aware feature mixing strategy that interpolates both sa...

Hong-Yang He, Xin-Yuan Song, Yan Zhong et al. · 1 citation
Preprint Sep 2026

Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification

Few-shot recognition with frozen visual features is especially fragile under domain shift and one-shot supervision, where a single labelled image is an unreliable estimate of its class. We ask how far this fragility can be reduced purely at test time, without retraining the encoder or augmenting the source domain. We p...

F. Rahman, S. Rohan, Mahmound Sayed et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.