WILDTRACE is introduced, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic.
Zixin Chen, Peng Liu, Haobo Li et al.· 0 citations
A-SR, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views, is proposed, a self-evolving agentic framework that shifts the control unit from expression edits to role-conditioned evidence views.
Wenxiao Zhao, Dong Liu, Kaiyi Xu et al.· 2 citations
SDABench is introduced, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics).
Chuhan Shi, Xiaoquan Ren, Sicheng Song et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.