Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par...
Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a cond...
UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and th...
Yi-Jia Fan, Zi-Qi Huang, Zhongang Cai et al.· 0 citations
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al.· 2 citations
This work introduces a novel prior-guided, multi-stage reconstruction pipeline that fuses geometric and motion priors to build a physically plausible and temporally stable foundation, and applies a universal trajectory-preserving technique, which safeguards high-quality motion by decoupling the appearance optimization...
Xiaosheng He, Feng-Lin Liu, Lin Gao et al.· IEEE Transactions on Visuali...· 0 citations
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 5 citations
Apple-PI is introduced, the first benchmark that anchors video-model evaluation explicitly in physical laws, and is positioned as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.
Runmao Yao, Kairui Hu, Yukang Cao et al.· 1 citation
Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.
Xiaoyang Han, Jianhua Li, Kewang Deng et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.