GeoAnchor decomposes 3D spatial information into three complementary components: position latents for object grounding, direction latents for relational orientation, and geometry latents for scene structure, and is introduced a collaborative training strategy that guides the model from local spatial perception to comprehensive 3D understanding.
Hao Li, Han Fang, Zixin Pan et al.· arXiv.org· 0 citations
A PETL-based framework named UniGR is presented that unifies granularity and reliability for robust and efficient TPR, and a multi-granularity relational adapter (MRA) is designed to capture both coarse-grained global and fine-grained local relational features among tokens.
Jingchen Hao, Jiang Liu, Zhen Peng et al.· Annual International ACM SIG...· 0 citations
NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation, and a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning.
Peng Cai, Zhaofan Zou, Shifa Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.