Skip to content

Author

Zhongang Cai

MMLab@NTU, Nanyang Technological University

We have 8 of 72 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Looped Diffusion Transformer

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par...

Yong Xien Chng, Tian-Yi Chen, Wen-Wen Tong et al. · 0 citations
#machine learning Preprint Sep 2026

One-Step Next-Latent Prediction Is Not a World Model

Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a cond...

Shi-Tong Wang, Zhongang Cai, Yu Hong · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and th...

Yi-Jia Fan, Zi-Qi Huang, Zhongang Cai et al. · 0 citations
Preprint Sep 2026

SenseNova-U1.5: Towards Native Unified Visual Intelligence

Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.

Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al. · 2 citations
Jul 2026

Consistent 4D Appearance Editing with Gaussian Splatting.

This work introduces a novel prior-guided, multi-stage reconstruction pipeline that fuses geometric and motion priors to build a physically plausible and temporally stable foundation, and applies a universal trajectory-preserving technique, which safeguards high-quality motion by decoupling the appearance optimization...

Xiaosheng He, Feng-Lin Liu, Lin Gao et al. · 0 citations
#machine learning Preprint Aug 2026

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.

Junhua Xu, Rui-Si Wang, Fanyi Pu et al. · 5 citations
Preprint Jul 2026

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Apple-PI is introduced, the first benchmark that anchors video-model evaluation explicitly in physical laws, and is positioned as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.

Runmao Yao, Kairui Hu, Yukang Cao et al. · 1 citation
Preprint Jul 2026

Vision as Unified Multimodal Generation

Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.

Xiaoyang Han, Jianhua Li, Kewang Deng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.