A novel method called DKA is presented, which introduces a limited set of concept-supervised data to enhance the knowledge base, effectively solving the reasoning shortcut problem and improving the applicability of the NeSy system.
Yu-Feng Li, Xiaowen Yang, Wenda Wei et al.· Proceedings of the 32nd ACM...· 0 citations
Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination. Most benchmarks either focus on single-agent task completion or summarize multi-agent behavior with overall task success rates, which can obscure coordination failures such as duplicated work, violations of ordering constraints, resource contention, and desynchronized handoffs. In this paper, we introduce CoCoBench, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks. CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively. We evaluate 11 leading MLLMs across different coordination modes, observation inputs, and numbers of agents. The results show that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types. These findings point to new directions for designing targeted model architectures and improving multi-agent coordination ability.
Yang Chen, Ye-Xin Xie, Li-Rong Che et al.· 0 citations
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.
Bo-wen Zhang, Junwei He, Wen Wang et al.· arXiv.org· 0 citations
Recent advancements in neuro-symbolic learning (NeSy) have shown significant promise in integrating deep learning with symbolic reasoning, offering both interpretability and generalization. However, the prevalence of reasoning shortcuts, where the NeSy system predicts incorrect intermediate concepts while maintaining high final accuracy, poses a substantial challenge. This is especially problematic in domains requiring reliable and transparent decision-making. Inspired by recent theories, we find that existing methods fail to address the reasoning shortcut issue when the knowledge base lacks sufficient complexity, highlighting their vulnerability in real-world applications. In this work, we present a novel method called DKA to address this issue. It introduces a limited set of concept-supervised data to enhance the knowledge base, effectively solving the reasoning shortcut problem and improving the applicability of the NeSy system. Theoretical analysis reveals that DKA can reduce shortcut risks with improved data efficiency. Empirical studies across multiple tasks within various neuro-symbolic frameworks also verify the effectiveness of the DKA method.
Yu-Feng Li, Xiaowen Yang, Wenda Wei et al.· Proceedings of the 32nd ACM...· 0 citations
CAST (Credit Assignment from Solver Teachers), which converts value changes in a game solver's state value into solver advantages and injects them into RLVR as turn-level signals and achieves the highest average zero-shot performance on ALFWorld and WebShop.
Yu Wang, Yi-Kai Zhang, Wentao Shi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.