Findings show that a skill document can serve as the starting point for a continuous control that improves both answers and actions without updating the backbone.
Xi-Jia Tao, Yi-Hua Teng, Xin-Yu Fu et al.· arXiv.org· 3 citations
The Noisy Test-time Reinforcement Learning framework (NTRL-Code) is proposed, which enables robust self-evolution of code LLMs using only unlabeled noisy data during the testing stage, and employs an abstract-syntax-tree (AST)-based structural aggregation mechanism to estimate a proxy target from multiple candidate pro...
Xi-Kai Yang, Hieu Trung Nguyen, Dun-Yuan Xu et al.· 0 citations
GraphDroid is proposed, an intent-driven GUI testing framework that integrates a cluster-based memory mechanism to effectively identify uncovered functionalities from historically visited states for comprehensive application testing and adopts an asynchronous intent generation paradigm that eliminates the latency bottl...
Xiao-Lei Li, Jia-Lun Cao, Zhijian Hou et al.· 0 citations
This work proposes Skill-Conditioned Gated Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation, and shows that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption.
Jiazhe Huang, Xiao Chen, Xiao Luo et al.· arXiv.org· 5 citations
FISA is proposed, a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases that generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion.
Chun-Yang Jiang, Pingping Zhang, Yuzhi Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.