Skip to content

Author

Yimao Cai

We have 4 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

Stage-Guided Per-Step Optimization (SGPO) is proposed for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 0 citations
Preprint Jul 2026

HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models Inference

HEMERA is presented, a heterogeneous memory-centric accelerator for efficient Mamba-2 inference that reformulates the matrix-form SSD computation into an algebraically equivalent streaming-recursive dataflow that avoids quadratic intermediate storage while preserving the original computation.

Hao Ding, Ling Liang, Ruitong Qiao et al. · 0 citations
Preprint Aug 2026

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

PAST is proposed, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty and establishes a dual adaptive coordination mechanism that balances the extrinsic and intrinsic rewards.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 0 citations
Review Jul 2026

Pixel-Space Diffusion Transformers

Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creates a mismatch between reconstruction and generation objectives. These limitations have renewed interest in pixel-space diffusion, which models raw pixels directly, removes the VAE bottleneck, and supports end-to-end optimization. This formulation better matches the demands of high-fidelity generation but introduces challenges in high-dimensional modeling, including noise scheduling, loss weighting, token efficiency, and scalable architecture design. Pixel-space modeling also offers a promising basis for unified multimodal systems: raw pixels, text, and task conditions can be represented in a shared token space and jointly processed by a single Transformer, narrowing the gap between visual understanding and generation. This paper reviews Pixel-Space Diffusion Transformers (pDiTs) from the perspectives of model architecture, continuous generative mechanisms, and unified multimodal modeling. We summarize representative methods, identify key technical challenges, and discuss future directions toward high-fidelity, end-to-end vision foundation models that integrate generation and understanding.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.