Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs...
Hao-Min Qi, Xiang-Zhe Xu, Yi-Ming Huang et al.· 0 citations
A search-based memory framework called BOOKMARKS for active grounding, which retains access to the full preceding storyline and collects task-relevant information on demand, and improves next-action fidelity in a five-dimensional evaluation.
This work proposes Process-Scorer Guided Adaptive Tree Rollout (PATR), a quality-aware rollout framework for multi-turn agent RL that uses task-appropriate process feedback to score partial trajectories, selectively branches from promising states, reuses shared prefixes, and conservatively stops degenerate paths to red...