Skip to content

Perceptual Flow Matching for Few-Step Generative Modeling

Jul 2026 · arXiv.org · Vol abs/2607.03524 · 2 citations · 39 references
Computer Science

TL;DR

Perceptual Flow Matching supervises flow matching in a perceptual feature space using pretrained perceptual models, which substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality.

Abstract

We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.

View source

Similar papers

Jul 2026

RFMSR: Residual Flow Matching for Image Super-Resolution

Residual Flow Matching for Image Super-Resolution (RFMSR) is proposed, a vision-only framework that centers the source distribution at the LQ latent, reducing transport distance and preserving structural priors throughout the flow trajectory.

Shuwei Huang, Tianyao Luo, Jicheng Liu et al. · 1 citation
Conference 2026

Shortcut Diffusion Training With Cumulative Consistency Loss: An Optimal Control View

This paper forms few-step generation as a controlled base generative process, and shows that self-consistency loss can be understood through the lens of optimal control, and draws a connection between this approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generation.

Paribesh Regmi, S. Ghimire, Rui Li · 0 citations
Preprint Aug 2026

Energy-Guided Flow Matching

Energy-Guided Flow Matching is introduced that explicitly models a coarse-to-fine generative trajectory by moving endpoint that evolves smoothly from low-frequency image to clean image and requires no adaptation of the backbone and training data.

Haoyang Tong, Yu He, Fang Li et al. · 0 citations
Preprint Aug 2026

MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution

This work proposes MeanSR, a one-step perceptual SR method that learns an LR-conditioned average velocity field to directly capture the finite-time transition from degraded or noisy inputs to plausible HR outputs and introduces a Stage-Aware Temporal Sampling strategy to improve trajectory learning.

Axi Niu, Jiawei Kou, Kang Zhang et al. · 0 citations
Preprint Sep 2026

Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations

Existing Conditional Flow Matching (CFM) formulations describe transport progress using an explicit interpolation coordinate, commonly interpreted as time, assuming that a single global variable adequately represents a sample's position along the generative trajectory. In restoration tasks, however, transport progress is sample-dependent because the initial distribution may exhibit varying statistical dependencies with the target distribution. Thus, samples at the same interpolation coordinate can differ substantially in degradation level, distance to the target distribution, and restoration difficulty. We investigate whether signal representations learned by discriminatively trained models provide a meaningful description of generative transport state in CFM-based restoration. Through systematic latent-space analysis, we show that discriminative representations organize according to degradation severity and follow a consistent trajectory toward the clean-data manifold during generation. Motivated by these observations, we introduce the Discriminative Flow-State Hypothesis, which posits that discriminative representations encode a transport state governing generative restoration. Based on this hypothesis, we propose Discriminative Flow Matching, which conditions the Flow-Matching velocity field on Discriminative Flow-State Representations rather than explicit time coordinates. Experiments on speech enhancement and image denoising show that these representations characterize restoration progress, enable adaptive inference, and consistently outperform CFM and diffusion-related baselines. Our findings suggest that discriminative representations provide an effective state-aware alternative to explicit time conditioning and offer a novel perspective on the relationship between discriminative and CFM-based generative modeling.

Unknown authors · 0 citations
Preprint Aug 2026

Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model

Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-matching objectives fit visual data distributions but do not explicitly penalize such physical violations. Existing remedies can improve physical consistency, but typically add substantial inference or training cost. Candidate-selection methods generate and score multiple videos, while gradient-based world-model guidance repeatedly decodes and re-encodes intermediate estimates. Generator-internal refinement adds perturbation and re-denoising loops, whereas post-training requires curated data and additional optimization. We propose Off-Manifold Refinement (OMR), an inference-time method that instead injects world-model feedback directly into a single sampling trajectory. During scheduled middle ODE steps, we augment the generator velocity with the gradient of an adapter-space V-JEPA 2.1 surprise energy. This external correction can move the latent away from the uncorrected sampling trajectory and toward regions ranked as more physically plausible by the frozen predictor, after which the generator continues rendering from the corrected state. A small trained latent-to-embedding adapter keeps the gradient tractable at inference, and both the video generator and the world model remain frozen. On our fixed 400-prompt VideoPhy-2 detailed subset, OMR lifts the joint Semantic-Adherence-and-Physical-Commonsense metric from 47.0% to 52.0% (+5.0pp absolute, +10.6% relative) over the base Wan2.2-T2V-A14B sampler. On a separate fixed 50-prompt efficiency subset, it requires $1.71 \times$ the base runtime rather than the multiplicative cost of reward/search alternatives. Project page: https://itruonghai.github.io/omr.

Hai Nguyen-Truong, Tuan-Anh Vu, Dang T. Huynh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.