Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par...
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al.· 2 citations
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 5 citations
PhysVICL-74 is introduced, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering and TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering.
Siyi Xie, Xuanke Shi, Jin-Sheng Quan et al.· 0 citations
Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.
Xiaoyang Han, Jianhua Li, Kewang Deng et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.