Conference
Jul 2026
3D vision-language question answering with explicit scene graphs and local topology priors
TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.
Kaixin Wu, Kunlin Zhou, Boxin Li et al.
· International Conference on... · 0 citations