#artificial intelligence
May 2026
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.
Kewei Xu, Xiaobe Lu, Shuofei Qiao et al.
· arXiv.org · 1 citation