Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachm...
Zhi-Yuan Qi, Jie-Rui Li, Yi-Fan Shen et al.· 0 citations
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for...
Tianjiao Yu, Xin-Zhuo Li, Yi-Fan Shen et al.· 0 citations
This work proposes ChronoVision, a multimodal framework designed to align visual logic with latent imagery, and introduces Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task.
This work proposes FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations, and efficiently infers pseudo frame-level alignments without human labeling.
Tongyan Wang, Zhengyuan Li, Muhan Lin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.