Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choi...
Hong-Yang Du, Yun-Fei Xie, Jun-Jie Ye et al.· 0 citations
Achieving accurate and efficient object pose estimation is a key goal in computer vision. Most existing methods rely on controlled environments, limiting their effectiveness in complex, dynamic, and unstructured real-world scenarios, especially for novel objects, severe occlusion, or sensor noise. Recent studies show t...
Hui Zhang, Yue Wang, Jianhao Jiao et al.· IEEE Transactions on Automat...· 0 citations
This work proposes SaLon3R, a novel framework for Structure-aware, Long-term 3DGS Reconstruction that effectively prunes the redundant 3DGS and resolves artifacts in a single feed-forward pass, and introduces a 3D Point Transformer to overcome geometric inconsistencies caused by long-term accumulative errors.
Jiaxin Guo, Tongfan Guan, Wen-Zhen Dong et al.· International Journal of Com...· 5 citations· ⚡1
A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.
Hui Zhang, Yue Wang, Kang An et al.· Signal, Image and Video Proc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.