Dyad is introduced, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions.
Yundaichuan Zhan, Wei-Shi Wang, Wen-Biao Liu et al.· 0 citations
Persona prompts ask language models to answer as particular kinds of people. We test whether relationships learned from these effects predict responses to new questions and remain useful across models and prompts. Across 57 attributes, three behavioral domains, and seven pairs of open 7 to 9B checkpoints, persona effec...
Yu-Fan Zhou, Yuxuan Liu, En-Ze Ma et al.· 0 citations
Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation analysis on current VLMs, evaluating both the comprehensiveness and suffic...
Xu-Rui Song, Wei-Shi Wang, Zhong-Qi Yue et al.· 2 citations
This work states that existing visual generative models are not yet ready for RL due to the following two fundamental drawbacks that undermine the foundations of RL.
Bohan Wang, Min Zhou, Zhongqi Yue et al.· Neural Information Processin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.