Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly s...
Ya-Kun Zhu, Yi Bin, Yu-Juan Ding et al.· 0 citations
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value project...
Yingcheng Liu, Tian-Yi Jiang, Yu-Juan Ding et al.· 0 citations
This survey establishes workload-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven-dimensional terminal competence profile, and provides a unified basis for studying terminal-mediated agency across software engineering and emerging application domains.
Yi Bin, Xiao-Yang Yuan, Hao Zeng et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.