Alignment-Guided Flow Transformer (AGFT) is presented, a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation.
Sheng-Chao Hu, Peng Wang, Qi-Yang Zhou et al.· 0 citations
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizin...
Sheng-Chao Hu, Peng Wang, Ji-Feng Hu et al.· 0 citations
Modular TTT is proposed, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions and finds that small learning-rate initialization, weight decay, and a single-layer nonlinea...
Bo-Hao Tang, Zhen Qin, Yu-Qi Pan et al.· 1 citation
FailForge is proposed, an agentic framework that converts failed rollouts into training signal, and recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline.
Dongyi Lv, E. Fushun, Aichen Cai et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.