Generative recommendation, has been attracted a surge of attentions in industrial and academic research community, towards to build more smart system to build next-generation recommender. Under the significant developing wave of large language model, our team have been developed Semantic ID based OneRec/OneRec-V2. Thes...
Jiang-Xia Cao, Hao Peng, Wen-Long Xu et al.· 0 citations
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shapes which trajectories are reachable. We study post-violation recovery admission, where p...
Chen-Yu Zhou, Qi-Liang Jiang, Shu-Ning Wu et al.· 1 citation
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction o...
Chen-Yu Zhou, Qi-Liang Jiang, Shu-Ning Wu et al.· 0 citations
PLCBENCH is presented, to the authors' knowledge, the first real-PLC hardware-in-the-loop (HIL) framework for characterizing this cyber-to-physical capability and its boundaries and it combines vendor-native interaction, commercial PLC execution, closed-loop reduced-order process simulation, and independent outcome ver...
Yitian Zhou, Jing-Yu Zheng, Qi-Liang Jiang et al.· 0 citations
Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers. Cert...
Chenyu Zhou, Qiliang Jiang, Shuning Wu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.