An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically e...
Peng Xia, Ru-Jun Han, Zifeng Wang et al.· 4 citations· ⚡1
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...
Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al.· 0 citations
This work designs five types of multimodal tasks across text, molecular SMILES strings and images, and curates the datasets, demonstrating the feasibility of unifying multiple cross-modal chemical tasks within a single foundation model and enabling more intuitive, visual human-AI interaction.
Qian Tan, Di Zhang, Ben Gao et al.· arXiv.org· 16 citations
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically e...
Peng Xia, Ru-Jun Han, Zifeng Wang et al.· 3 citations· ⚡1
Environment Harness is proposed, a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic, enabling continuous, targeted co-evolution of the policy and its environment.
Chengsong Huang, Zifeng Wang, Ru-Jun Han et al.· 11 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.