Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three...
Jun-Yu Guo, Shangding Gu, Ming Jin et al.· 0 citations
TRACE is presented, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning r...
Wen-Jun Xiong, Shengtao Zhang, Shangding Gu et al.· 0 citations
The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.
X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al.· 1 citation
Memory-Aware Propagation and Link Enforcement Guard, MAPLE-Guard, a memory-link guard for memory-enabled MAS, suggests that memory-aware link enforcement covers a gap left by prompt-level and topology-level defenses.
Wen-Jun Xiong, Yi-Jin Zhou, Jia-Qian Wang et al.· 0 citations
SafeClawArena is developed, a benchmark of 406 adversarial tasks executed in containerized replicas of real agent platforms with canary-marked credentials and evaluated via automated taint tracking across nine output channels, exposing the inadequacy of current defenses and suggesting directions for future hardening.