Lens3D: Target-Conditioned Visual Foveation for Fine-Grained 3D Understanding
Experiments demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance, and establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.
Jun-Ming Huang, Zi-Ni Chen, Shuai-Ying Hou et al.
· 0 citations