Skip to content

Author

Qi-Feng Chen

We have 10 of 205 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

HelixWorld: A Real-time Interactive Audio-Visual World Model

World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive aud...

Lei Ke, Jia-Hao Pan, Ze-Yue Tian et al. · 0 citations
Review Sep 2026

The Past Frames the Future: Memory for Autoregressive Video Generation

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...

Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al. · 0 citations
Preprint Aug 2026

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

USR-Drive is proposed, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation and delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the...

Li-Heng Chen, Haokai Pang, Cheng Su et al. · 0 citations
Jul 2026

IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

This work proposes IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations and constructs Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence.

Jin-Jian Wu, Jiaqi Tang, Wei Wei et al. · 0 citations
Preprint Aug 2026

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

Remember-R1 is proposed, a reinforcement learning framework that mitigates long-context visual forgetting by applying process-level supervision directly on the original reasoning trajectory, demonstrating its effectiveness in mitigating long-context visual forgetting.

Jianmin Chen, Jiaqi Tang, Wei Wei et al. · 0 citations
Preprint Aug 2026

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editi...

Zhe-Fan Rao, Bin-Yi Zou, Xuanhua He et al. · 0 citations
Preprint Aug 2026

MSEditor: Toward Consistent Multi-Shot Video Editing

MSEditor is proposed, the first framework designed specifically for consistent multi-shot video editing, which significantly outperforms existing methods on the authors' curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.

Kun-Yu Feng, Yue Ma, Bing-Yuan Wang et al. · 1 citation
Jul 2026

Data Pyramid for Embodied Manipulation

This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...

Yifan Ye, Yankai Fu, Ya-hui Lv et al. · 4 citations
Jul 2026

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

LeapBot-WA establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor and introduces the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift.

Pei Liu, Nan Zheng, Lang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.