Multimodal large language models (MLLMs) are increasingly used as automated judges for instruction-based image editing and as reward signals for model training. However, systematically auditing whether these judges are influenced by cues irrelevant to editing quality is challenging because visual interventions may them...
J-Access is proposed, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway, positioning J-Access as a model-level diagnostic for assessing residual susceptibility in unlearn...
Zirui Song, Hua-Xing Liu, Xiang Wang et al.· 1 citation
SocialMaze is introduced, a benchmark that organizes six tasks across social deduction games, daily-life interactions, and digital community platforms along three descriptive design axes: deep reasoning, dynamic interaction, and information uncertainty.
Zi-Xiang Xu, Yan-Bo Wang, Yue Huang et al.· 1 citation
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operatio...
Zixiang Xu, Sixian Li, Hua-Xing Liu et al.· arXiv.org· 2 citations
Inspired by the Big Five model's dimensional view of personality, a framework that reframes authorial writing characteristics as coordinates within a unified and interpretable space is proposed, which improves authorial expressiveness while preserving semantic fidelity.
Jinghui Zhang, Lang Gao, Ao Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.