DocOps is introduced, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities that exposes the capability boundaries of agents in maintaining global document consistency.
Jiazhen Jiang, Boxi Cao, Lingyong Yan et al.· arXiv.org· 1 citation
DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform, is introduced, showing that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework.
Zechun Niu, Yu-Kun Zhao, Jia-Xin Zhang et al.· 0 citations
A lightweight training framework that learns a single Behavior-Equivalent Token that substantially reduces inference cost and frees nearly the entire context window for user inputs and model outputs.
Jiancheng Dong, Pengyue Jia, Jingyu Peng et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.