This work proposes Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding, which outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.
Haojie Huang, Xin-Lei Yu, Cheng-Ming Xu et al.· 3 citations
Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent capabilities into image generation, but they either prescrib...
Jiahao Zhao, Xiao-Min Yu, ZhongXiang Sun et al.· 3 citations
Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not...
Wen-Bin Wu, Yu-Zhong Wu, Yu-Fan Xu et al.· 0 citations
LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.
Hongyu Qu, Jianzhe Gao, Xiaobin Hu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.