Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired kn...
Rui-Fan Zuo, Guo-Cheng Hu, Wan-Shui Gan et al.· 0 citations
Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts in...
Manwen Liao, Xinyu Lian, Jian Mao et al.· 0 citations
OccAnyScene is proposed, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model which employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple...
Junjie Liu, Wan-Shui Gan, Zi-Tong Dai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.