LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is diffi...
Hao-Ran Li, Z. Ge, Xiao-Ming Yuan et al.· 0 citations
This work proposes SCoRE (Selection and Consolidation for Robust Evidence), a unified agent loop for explicit evidence selection and consolidation, which decouples final reasoning from exploratory trial-and-error while ensuring strict visual grounding via indexed claim-to-image linkages.
Yucheng Shen, Lingyong Yan, Jiulong Wu et al.· 0 citations
Janus is introduced, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators to address label scarcity and extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.
Xi-Meng Liu, Qianlong Wang, Ying-Ming Mao et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.